@elmelis

Asst. Prof. @hseas @KempnerInst || Researcher @MSRNE || ML + NLP || Previously: @MIT_CSAIL NYU @IBMResearch @ITAM_mx

Cambridge, MA
Joined August 2010
David Alvarez Melis retweeted
4B parameters. ZERO distillation. 61.5% on SWE-bench Verified 🤯 Meet FrogNano 🐸: Qwen3.5-4B post-trained purely with RL on synthetic tasks from TaskPilot. Just 5 iterations × 300 tasks. Who said coding agents have to be huge? 🐸
55
115
26
1,180
117,328
Very nice project led by @sunnytqin and @kimiahmdh! Classical scaling laws implicitly assumed data was free, but nothing in life is free😆. So the right question is what is the *exchange rate* between fresh and derived tokens, and our results show that rate is far from constant.
(1/N) 🧵 Chinchilla assumes you'll never run out of fresh data. That era is ending! Compute keeps growing exponentially, but high-quality tokens don't. So what's the exchange rate between extra compute and fresh, high-quality data? We propose Compute-Data (CD) scaling laws, which measure the exchange rate between extra compute and fresh data.
1
3
27
4,147
My favorite finding from the paper: the "4 epochs is fine" rule-of-thumb really only holds near 3B params and 1× Chinchilla. Outside that regime, both the optimal number of epochs and whether paraphrasing helps at all shift a lot. More cool findings in Sunny's thread above.
4
262
David Alvarez Melis retweeted
We have an amazing lineup of invited speakers joining us in NYC this October, including @elmelis @mariannearr @Clement_Bonet_ @sitanch @YongxinChen1 @cdomingoenrich @ArthurGretton @yjelid @k_neklyudov @ssahoo_ @SchiffYair @sherryyangML @SoojungYang2
📣 We are excited to announce the workshop Emerging Directions in Probabilistic Modeling: Methods and Scientific Applications 📅 Oct. 5–7, 2026 📍 Flatiron Institute & IBM Research, NYC Applications are welcome by Aug. 22: edpmworkshop.github.io
6
1
31
5,382
Very excited about this work, led brilliantly by @kimiahmdh Does your data mixture give you synergy vibes? Like math-and-code-kind-of-vibes? Here's one (technically, two) ways to turn those vibes into concrete estimates, which can then be used to optimize mixture design.
Can we tell whether data domains cooperate or compete during pretraining? Adding code to the mix makes models better at math, while some other combinations hurt each other. We call this data synergy. Turns out you can incorporate data synergy into scaling laws and estimate it 🧵
7
39
7,903
David Alvarez Melis retweeted
We don’t always know what problems are hard for LLMs. So devs evaluate on tasks HUMANS find hard or on broad benchmarks. What if we could instead anticipate which scenarios a model will fail on—all without evaluating specific input examples? 🧵NEW PAPER by @jenniferlumeng &al
8
48
6
276
35,027
This was a fun one! And a real treat to be a (small) part of it. Partly because the paper formalizes a bunch of things about scale/data that felt plausible but fuzzy to me before, and partly because watching Ekdeep in action is a treat of its own. He's one of a kind.
We take for granted that larger models are better than smaller ones, but why is this so? Our new paper, led by Jing Huang and @EkdeepL, traces this to a data-induced competition for resources (neurons), using formal analysis, idealized tasks, and real pretraining.
1
2
22
3,229
It also suggests that the classic (learning theory) way to think about model capacity in isolation misses an important part of the story. We ought to think about capacity 𝘳𝘦𝘭𝘢𝘵𝘪𝘷𝘦 𝘵𝘰 task diversity in the training data.
1
2
163
Plenty of open questions left about how to choose mixtures for a given scale, transfer to real pretraining, post-training as a separate axes, etc, but this paper lays solid foundations to think about all of these. Link: arxiv.org/abs/2605.29548
2
164
David Alvarez Melis retweeted
We have a last-minute internship opening for summer/fall 2026---working with Samy Jelassi, who just joined us! Apply here: apply.careers.microsoft.com/…
14
35
2
496
63,620
Our Data-Centric ML group is at ICLR 🇧🇷this week. I couldn't make it this year 😰, but @SaraKangaslahti, @JonathanGeuter, @rach_it_ are there. Find them, say hi. Quick rundown 👇
1
5
23
1,561
Monday (SPOT Workshop): RL Excursions during Pre-training: how early is too early for on-policy learning? TL;DR: a look at when on-policy RL starts helping (or hurting) during pre-training. w/ @rach_it_, @clara_mohri , @sunnytqin, @ShamKakade6. rl-excursions.github.io
1
2
203
David Alvarez Melis retweeted
Teach your own model to manage its context with Memento. We're open sourcing everything. Brought to you by AI Frontiers, a boutique lab within Microsoft Research
9
33
6
444
59,045