@elmelisi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Asst. Prof. @hseas @KempnerInst || Researcher @MSRNE || ML + NLP || Previously: @MIT_CSAIL NYU @IBMResearch @ITAM_mx
Cambridge, MA
Joined August 2010
- Tweets856
- Following2.3K
- Followers2.1K
- Likes2.5K
David Alvarez Melis retweeted
4B parameters. ZERO distillation. 61.5% on SWE-bench Verified 🤯
Meet FrogNano 🐸: Qwen3.5-4B post-trained purely with RL on synthetic tasks from TaskPilot.
Just 5 iterations × 300 tasks.
Who said coding agents have to be huge? 🐸
Very nice project led by @sunnytqin and @kimiahmdh!
Classical scaling laws implicitly assumed data was free, but nothing in life is free😆. So the right question is what is the *exchange rate* between fresh and derived tokens, and our results show that rate is far from constant.
(1/N) 🧵 Chinchilla assumes you'll never run out of fresh data. That era is ending!
Compute keeps growing exponentially, but high-quality tokens don't. So what's the exchange rate between extra compute and fresh, high-quality data?
We propose Compute-Data (CD) scaling laws, which measure the exchange rate between extra compute and fresh data.
David Alvarez Melis retweeted
We have an amazing lineup of invited speakers joining us in NYC this October, including
@elmelis @mariannearr @Clement_Bonet_ @sitanch @YongxinChen1 @cdomingoenrich @ArthurGretton @yjelid @k_neklyudov @ssahoo_ @SchiffYair @sherryyangML @SoojungYang2
📣 We are excited to announce the workshop
Emerging Directions in Probabilistic Modeling: Methods and Scientific Applications
📅 Oct. 5–7, 2026
📍 Flatiron Institute & IBM Research, NYC
Applications are welcome by Aug. 22:
edpmworkshop.github.io
Very excited about this work, led brilliantly by @kimiahmdh
Does your data mixture give you synergy vibes? Like math-and-code-kind-of-vibes? Here's one (technically, two) ways to turn those vibes into concrete estimates, which can then be used to optimize mixture design.
David Alvarez Melis retweeted
We don’t always know what problems are hard for LLMs. So devs evaluate on tasks HUMANS find hard or on broad benchmarks. What if we could instead anticipate which scenarios a model will fail on—all without evaluating specific input examples?
🧵NEW PAPER by @jenniferlumeng &al
This was a fun one! And a real treat to be a (small) part of it. Partly because the paper formalizes a bunch of things about scale/data that felt plausible but fuzzy to me before, and partly because watching Ekdeep in action is a treat of its own. He's one of a kind.
We take for granted that larger models are better than smaller ones, but why is this so? Our new paper, led by Jing Huang and @EkdeepL, traces this to a data-induced competition for resources (neurons), using formal analysis, idealized tasks, and real pretraining.
ALT Title card for a research paper. The title reads "Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention." Authors listed: Jing Huang, Daniel Wurgaft, Rachit Bansal, Laura Ruis, Naomi Saphra, David Alvarez-Melis, Andrew Lampinen, Christopher Potts, and Ekdeep Singh Lubana. A Goodfire logo appears below the names. Author affiliations: Stanford University, Kempner Institute at Harvard University, MIT, and Anthropic.
It also suggests that the classic (learning theory) way to think about model capacity in isolation misses an important part of the story. We ought to think about capacity 𝘳𝘦𝘭𝘢𝘵𝘪𝘷𝘦 𝘵𝘰 task diversity in the training data.
Plenty of open questions left about how to choose mixtures for a given scale, transfer to real pretraining, post-training as a separate axes, etc, but this paper lays solid foundations to think about all of these. Link: arxiv.org/abs/2605.29548
David Alvarez Melis retweeted
We have a last-minute internship opening for summer/fall 2026---working with Samy Jelassi, who just joined us! Apply here: apply.careers.microsoft.com/…
Our Data-Centric ML group is at ICLR 🇧🇷this week. I couldn't make it this year 😰, but @SaraKangaslahti, @JonathanGeuter, @rach_it_ are there. Find them, say hi. Quick rundown 👇
Sunday (Multimodal Intelligence Workshop): Stop Training for the Worst. TL;DR: progressive unmasking accelerates masked diffusion training by, well, not training for the worst case upfront. w/ @Jaeyeon_Kim_0, @JonathanGeuter, @ShamKakade6 , @sitanch. arxiv.org/abs/2602.10314
Monday (SPOT Workshop): RL Excursions during Pre-training: how early is too early for on-policy learning? TL;DR: a look at when on-policy RL starts helping (or hurting) during pre-training. w/ @rach_it_, @clara_mohri , @sunnytqin, @ShamKakade6. rl-excursions.github.io