@Mlbot4

Exploring knowledge | Machine Learning Engineer

Joined March 2019
Gassel retweeted
Why is RL so sensitive to training-inference mismatch (TIM)? With each training step, the trainer is pulled toward the sampler, resulting in accumulating drift. We exploit this intuition to devise a novel correction method to stabilize RL under TIM, called Score Centering.
16
52
8
474
50,059
"DiffusionGemma as Jev" showcases the power of non-autoregressive architectures. While Jev demonstrates the value of rapid decision models, running DiffusionGemma in this paradigm leverages canvas diffusion to evaluate structured choices in a single parallel pass: ⚡ ️Massive Parallelism: Denoises across an open canvas in a single step instead of sequential autoregressive token generation (~0.2s on a DGX spark). 🧠 Full Bidirectional Attention: Allows every option to attend to the full context concurrently, yielding well-calibrated decision distributions. 👁️ Multimodal Grounding: Inherits Gemma 4's spatial vision capabilities for complex visual and text decisions. Read more about this approach here: github.com/vllm-project/vllm… nitter.cf/mmastrac/status/210037… nitter.cf/mmastrac/status/210062…
I ran some real, live evals on Jev vs DiffusionGemma-as-Jev (my patch for vLLM!) DiffusionGemma comes out as the winner, I think. Headlines: Is Jev faster than DiffusionGemma? No ❌ (API vs DGX Spark) Is Jev smarter than DiffusionGemma? No ❌ (they're roughly tied!)
36
272
47
2,667
214,744
Gassel retweeted
Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand unfamiliar data, develop ML and data analysis pipelines and produce results despite limited data, unspecified goals and/or very limited feedback. 1/8
WeirdML v2 is now out! The update includes a bunch of new tasks (now 19 tasks total, up from 6), and results from all the latest models. We now also track api costs and other metadata which give more insight into the different models. The new results are shown in these two figures. The first one shows an overview of the overall results as well as the results on individual tasks, in addition to various metadata. The second figure shows cost vs performance and shows a clear scaling with better results for higher costs. We also have a very varied pareto frontier with 11 models from 6 different companies having the best accuracy for a given cost for at least some of the cost range. Grok 3, Claude Opus 4 and GPT 4.5 are the ones that underperform for their costs, while Gemini pro and o3 pro have the best results at the highest costs. Qwen3 30B3A, grok 3 mini and deepseek R1 also each represent a good chunk of the pareto frontier.
51
53
13
617
89,612
Today’s LLMs still write like typewriters: one token at a time. This sequential process creates a hard inference bottleneck. We're introducing Uno, a diffusion-augmented LLM that delivers autoregressive quality at diffusion speed. It’s a lossless speedup method that accelerates generation without degrading response quality. With Uno, K2-Horizon-7B outperforms state-of-the-art diffusion methods in both quality and throughput, delivering up to a 2.2× speedup with no loss in quality. Paper: arxiv.org/abs/2609.04010 Model available at: huggingface.co/IFM/K2-Horizo…
82
177
139
1,022
222,019
PrismML team fit a 27B model into 5.9 GB. We uncensored it without changing a single weight. 🐳 Introducing OrcaRouter Ternary Bonsai 2 27B Uncensored. Traditional abliteration modifies weights. On a ~1.72-bit ternary model, that means re-quantization — potentially destroying the quality preserved by QAT. So we moved abliteration into the runtime. → 0 weights modified → 0 re-quantization → original 5.9 GB pack stays bit-identical → 129 residual intervention sites → adjustable at inference → runs locally on Apple Silicon No modified checkpoint. Bring the original Bonsai pack + our runtime. Open source: github.com/Continuum-AI-Corp…
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
88
273
44
3,903
301,673
Gassel retweeted
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
613
1,515
737
15,140
4,981,795
Replying to @AlberFuen
We have jev at home. The arch is good for RL with state-dependent actionsets, so here is the same arch playing doom and chess (shittily) with two different controllers.
1
2
29
2,175
I reverse-engineered a jev-like architecture given its type. You can find the repo here to train your own jevlikes: github.com/vinnylarouge/jevl…
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
48
168
22
2,018
164,324
I have conducted an audit of Anthropic's finances. What I have found is so shocking that I am calling for a Congressional investigation. Anthropic is not just seeking regulatory capture. It has built a regulatory capture machine that cannot be turned off. Structural financial incentives make it impossible for Anthropic -- I call it the Anthropic Network -- to turn off its own AI doom cycle. It starts with METR. Dario Amodei proposes "third-party evaluators" to assess the risk of Anthropic's models. He proposes METR for this purpose. But METR is financially dependent on the Anthropic's success -- specifically, on the explosive growth of more than $7 billion dollars in Anthropic stock. Dustin Moskovitz invested this stock into Good Ventures Foundation, where it represents the majority of that organization's portfolio. And GVF is the overwhelming funder of the entire Anthropic Network ecosystem. This stock was worth $500 million early last year. It is worth more than $7.7 billion just ~16 months later. METR -- and all of those building a career its parent organizations -- cannot afford to disrupt that growth. Because if Anthropic goes under, many of the organizations that fund METR go under as well. But if Anthropic succeeds, METR and its parent organizations become more richly financed to regulate AI -- something those at METR want very much. The "third-party evaluator" is not "third-party" at all. The evaluator is on Anthropic's payroll. If this were the end of it, that's bad. But that isn't all. The same organizations that fund METR also fund the many organizations, such as the Tarbell Center, that promote AI Doom. The Tarbell Center publishes AI Doom articles in The Verge, Science, LA Times, The Dispatch, TIME, and others. They are selling the problem, and then selling the solution to the problem -- from the same money pile: Anthropic's. All of these organizations are financially dependent on the same exploding $7 billion money pile. As Anthropic grows more and more powerful, its AI Doom Machine grows better and better financed -- louder and louder. Meanwhile, the regulatory regime seeded in METR grows larger to solve the increasingly loud -- now hysterical -- problem of AI Doom that the Anthropic Network itself created. From this standpoint, as Anthropic becomes more powerful, AI might be getting scarier, sure -- but the positive feedback loop also becomes more deafening -- independent of objective facts. This itself is an objective fact. The deafening AI Doom is part of an business model, that, as it expands, so too does the AI Doom messaging -- there is simply more money to do it. But the problem also goes in the other direction: If Anthropic dies, the Regulatory Regime and the AI Doom Machine are crippled or die. Neither METR nor Tarbell nor the other organizations in the Anthropic Network can allow that to happen. Hence, neither METR or the AI Doom Machine can be trusted to provide independent assessments of Anthropic's models or AI more broadly. They simply are not organizations independent of Anthropic. And Anthropic cannot detach itself from METR or Tarbell or countless other safety orgs (not shown here), either, because they drive hype for the models and the possibility of eventual regulatory capture, and Anthropic will not give that up willingly. What's more, the people at all of these organizations are all the same ecosystem, the same community. They just shuffle between organizations. The Anthropic Network is therefore, so long as it is successful, locked into a self-amplifying feedback loop inside an ideological monoculture. And that feedback loop is winning. That's what Jacob Coxon is. China is keeping messaging tight. That is why optimism for AI is so high in China. America has Anthropic: a massive company pushing anti-AI propaganda at a state level. Anthropic will either create hysteria until American AI slows down and China wins, or it will create fractures throughout American society with severe political consequences. Ironically, because of the structural financial incentives underpinning the Anthropic Network, it has become the same kind of self-amplifying virus that it fantasizes AI to become in the future -- while hiding its tracks just as carefully. It is the mirror of the same AI virus that it hypothesizes to consume America. Anthropic's business model, models itself after the very thing it claims to fear. Except Anthropic's ideology infects humans, not computers. Congress must investigate. Evidence and Github in next post. Then some supplementary figures.
1,555
9,631
2,042
34,001
6,473,833
Superintelligence should learn from experience through RL. Introducing FlashREINFORCE: Critic-Free, Single-Rollout, Asynchronous RL for Agentic Language Models Reinforcement Learning Should Do REINFORCE! github.com/yifanzhang-pro/Fl…
FlashREINFORCE: to our knowledge, the first open-source critic-free, single-rollout async LLM RL with 6,000+ stable updates. One-Batch REINFORCE + Sequence Trust Region + Sample-Mean Optimization. Paper: researchgate.net/publication… Code: github.com/NVIDIA-NeMo/labs-…
19
68
6
721
1,251,417
TL;DR: 1. Record pre-sft loss L0 of every example 2. During SFT, exclude the most-improved-compared-to-L0 examples in the batch from bprop (eg sort and slice or mul by zero) 3. Get a better pass@k for k>1 starting point for RL Seems simple and intuitive!
RL is expensive, so every step should count. ~1 yr ago, we showed xent SFT isn’t the best way to prepare for RL (arxiv.org/abs/2510.15020). Now, we propose TailSFT (arxiv.org/abs/2608.25756), a lightweight + principled way to directly improve coverage and get better post-RL perf.
21
62
7
740
88,232
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
1,042
3,086
1,720
28,406
6,768,444
DeepSeek releases DeepSeek-V4.1-Flash!! 🔥 This model is wild... the benchmarks are showcasing it's GPT-5.6 Sol level, yet it's only 552B params?! This feels like it shouldn't be possible, what's the catch? Perhaps benchmarkmaxxed? Let's look into the architecture and training of the model: The key focus seems to be on more aggressive KV cache compression: "We adopt a Causal Encoder-Decoder (CED) architecture, in which decoder global KV is projected from the final encoder hidden states. This design enables the model to activate 8B parameters per token during prefill and 16B during decode, which is particularly cost-effective for input-heavy agentic scenarios." "DeepSeek-V4 can be viewed as an SWA-based local-processing backbone augmented with compressed global context." The idea behind CED is that the decoder's KV cache is constructed from the encoder output, bypassing full decoder computation. They also introduce Compressed Sparse Attention 2 (CSA2) that has three operating modes that differ in how they obtain main KV, indexer K, and Top-K indices. (frankly I don't understand this part very well 😭) DeepSeek-V4.1-Flash uses a variant of mHC called single-pass mHC, and they also incorporate a 196B Engram module to decouple memorization from computation. The model is natively multimodal: a vision embeddings generated from "DeepSeek-ViT" (a pretty standard ViT arch) are passed jointly with the text tokens into the model. This is trained first with SigLIP loss then with autoregressive loss for the combined vision encoder+LLM. "we train DeepSeek-V4.1-Flash on a large-scale multimodal corpus comprising 45T tokens." They use FP4 KV cache with quantization-aware training to further save storage. Regarding post-training: "In this release, we refrain from introducing novel post-training algorithms." "at the current stage, the marginal return of engineering the data and environment pipeline substantially exceeds that of algorithmic novelty in post-training." They utilize the model itself to construct its own training environments, based on data they are getting from model use internally. "As we transitioned from DeepSeek-V3 to V4, the rapidly growing number and diversity of agentic training environments motivated us to build DeepSeek Elastic Compute (DSec), a production-grade sandbox platform for large-scale agentic training and evaluation." "We therefore introduce a scalar effort level 𝑏 as an explicit conditioning signal during reinforcement-learning." max --> b=100, high --> b=75, low --> b=50. "As the last stage of post-training, the final full-vocabulary OPD task is trained on datasets from all domains using over 40 teacher models." Damn, this is a dense report, I've barely touched the surface tbh, very interesting!! model: huggingface.co/deepseek-ai/D… paper: huggingface.co/deepseek-ai/D…
17
52
3
405
61,991
New open source music generator, YuE2 Even beats Mureka v9 and Minimax 3 on various benchmarks map-yue2.github.io/
77
237
55
2,460
424,662
A 35B language model running on an iPhone using only 1–2.5 GB of peak memory. No cloud. No remote server. No desktop GPU. Today, we’re open-sourcing Edge0 — a framework for running large AI models fully on-device.
394
836
149
10,277
715,947
This post looks like the start of a VERY sophisticated and well-funded PR operation to get support for Democrats to regulate AI into oblivion. Let me show you how it works: 1.) This guy, with minimal followers and no previous account activity, goes to the Wall Street Journal which publishes an exclusive with quotes from him on his resignation 18 minutes BEFORE this post goes up. Planning was clearly done in advance. 2.) Within hours, it has tens of thousands of reposts and the account has 100k+ followers. The post is punchy, quotable, it almost seems professionally written. The first three accounts to quote tweet it all do so within 15 minutes of the initial posting. Remember, this account had basically zero engagement beforehand, so an organic reach explanation seems unlikely. According to Grok those accounts are @_NathanCalvin (General Counsel at Encode AI), @peterwildeford (Head of Policy at the AI Policy Network), and @DKokotajlo (Head of the AI Futures Project), all of which are up-and-coming AI-Doomer policy advocacy nonprofits. The AI Futures Project website says it is funded “primarily” by the Survival and Flourishing Fund, which says on its own website that it has advised Jaan Tallinn, Skype creator and one of the leading investors in Anthropic, to grant over $2.5 million to the AI Futures Project since 2024. Encode AI says on its website that it is ALSO funded by the Survival and Flourishing Fund, which in turn says that it told Anthropic investor Jaan Tallinn to grant $516,000 to Encode AI in 2025. And wouldn’t you know it, the Survival and Flourishing Fund ALSO says it told Jaan Tallinn to grant $2 million to the AI Policy Institute, the 501(c)(3) affiliate of the AI Policy Network, as well. What are the odds that the first three quote tweets of Coxon’s post would all be major AI-restriction policy advocates funded generously by the same donor, who also happens to be one of the leading investors in, and a board member of, Anthropic, the company Coxon was resigning from? And all within 15 minutes of posting (two within ten)? 3.) Jacob Coxon doesn’t have much of a resume, but we do know that, in 2022, he got a $20,159 scholarship for the “long term future scholarship program” from the Good Ventures Foundation, one of the philanthropic vehicles of Dustin Moskovitz, a notorious AI-doomer who has spent tens if not hundreds of millions on policy advocacy to strictly regulate AI, while also being an Anthropic Investor himself. It also just so happens that the 14th person to quote Coxon’s post was @MaxNadeau_ (27 minutes after posting) who is the program officer for the Technical AI Safety team at Coefficient Giving, another of Moskovitz’s philanthropic spending vehicles. Max is not a frequent poster, his last posts before quoting Coxon were before Labor Day, but he was remarkably quick off the mark for this one. 4.) Basically every major Democrat politician and candidate has suddenly glommed on to this post, and conveniently, as the people cry out foe answers, Bernie Sanders already has a bill written to “ban super intelligence” and regulate AI into oblivion, and will be releasing later this week. The bill, among many other things, will create “a new cabinet-level federal agency to safeguard the public from the dangers of artificial intelligence” that will be “advised by an Artificial Intelligence Advisory Board comprised of experts on artificial intelligence.” Do you think, perhaps, Anthropic and its many investors who fund AI policy advocacy might have interest in getting to place a pet “expert” on the board of an entity that dictates what AI is and isn’t allowed to do? And isn’t it fortuitous that this whistleblower came forward with his oh-so scary stories so close in proximity to the release of the most radical piece of AI legislation ever introduced?
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
1,841
6,764
2,238
32,292
7,014,561
Check out our new work on Tail-Likelihood Reinforcement Learning (TailRL), extending maximum-likelihood RL from binary to continuous rewards. arxiv.org/abs/2609.02987 Rather than optimizing only mean reward, TailRL maximizes the expected log of upper-tail probabilities, naturally placing more weight on rare, high-reward rollouts. Its gradient can also be interpreted as a mixture of Best-of-(k) gradients. TailRL requires only a simple modification to the advantage function, making it easy to integrate into existing RL pipelines. Across object localization, maze navigation, GUI grounding, and code optimization, TailRL effectively exploits rare high-reward samples and scales better with increased inference-time sampling. Check out a detailed thread by @stablegradients.
Is RL optimizing the right objective? 🤔 Should we maximize mean reward? Best-of-k? Which k? Standard RL pulls on the mean and often the distribution collapses to a spike. The tail dies 🥲 We introduce Tail-Likelihood Reinforcement Learning (TailRL). It maximizes the mean reward while simultaneously maximizing coverage over high reward outputs. 🧵 1/n
8
44
3
428
66,627
MAI-Transcribe-2: the highest quality, cheapest transcription at the fastest speed! 10x faster that GPT-Transcribe. Now available on Microsoft Foundry.
27
57
28
602
44,527
today, we're releasing the largest open-source human audio preferences dataset, focused on the customer support use-case - 300K+ annotations by real people - 15 SOTA TTS models ranked (Sonic 3.6, Grok TTS, Simba 3.2, Eleven Labs v3) - 8 categories (IVR menus, empathy, escalations, refunds etc) dataset + benchmark + frontier plot below:
14
39
4
417
30,397
Breeze-TTS-2 is here. 🎙️ Real-time voice cloning, design, and direction in one model. 🤖 modelscope.ai/models/BreezeB… 🏆 #1 open-weight TTS on Artificial Analysis, plus first place on voice design, instruction following, and latency benchmarks. 🎭 Clone a voice while preserving timbre, rhythm, emotion, and style, or design and direct new Chinese and English voices with natural-language prompts and inline vocal events. ⚡ Under 40 ms TTFA and 3.1× real-time streaming on the warmed-up fast path. Eager inference uses ~7.7 GiB, with a 12 GB GPU recommended. 📜 Code: Apache 2.0. Weights, derivatives, and self-hosted outputs: research and non-commercial use only.
7
25
5
253
16,235