@tydsh

Co-founder of @Recursive_SI. ex-Meta FAIR Director. ex-Google. Reasoning, Optimization and Understanding LLM. Novelist in spare time. PhD in @CMU_Robotics.

California, USA
Joined December 2009
Early results from Recursive 🚀🚀 SotA results from our open-ended knowledge discovery system: 1️⃣NanoChat 5min pre-training (0.9372 bpb -> 0.9109 bpb, 2.8% lower Bits-Per-Byte than long-standing community SoTA) 2️⃣NanoGPT SpeedRun (79.7s -> 77.5s, 2.8% faster than long-standing community SoTA) 3️⃣GPU kernel optimization (Overall 7.8% better than SoTA performance in SOL- ExecBench, hosted by NVIDIA) To achieve that, our system automatically finds and combines innovations together to create better solutions than current ones carefully designed by expert humans in various domains. We have open-sourced resulting artifacts found by our system so you can check the output yourself. See a full breakdown and technical writeup: recursive.com/articles/first…
9
37
5
371
73,475
We got exact the same comments when we published Coconut😀. However, even with explicit CoTs, the model can still learn to hide the true thinking process. The ultimate solution is always to understand how the model works.
new: OpenAI & others quietly using loop transformers that don't show their 'thinking' when scaled up a leap forward on performance, but sparking concerns inside & outside OpenAI re: security as this takes off
9
3
1
104
14,131
😅 so our StreamingLLM (the original attention sink paper) strikes back again? 😅 -- Caveat: it is in post-training setting.
4
4
132
23,391
Interesting findings by our automatic system😆 It finds that some FlashInfer kernels have numerical issues, provides a PR and accepted by human experts🤓
A fun small win for automated research: We found and helped fix some edge cases that could affect inference performance in vLLM and SGLang. In our last blog post, we described a reward hacking judge we developed for performance optimization tasks. While applying the judge to some new autoresearch work, it discovered that some FlashInfer (a library underpinning vLLM and SGLang) kernels used a hard coded value of -50,000 as a masked-attention sentinel -- even though valid QK values can be smaller. Corner cases like this that are numerically wrong but silent can cause a huge amount of headache to find and fix, e.g., the historical debate around flash attention (arxiv.org/abs/2405.02803). Great use case for AI. github.com/flashinfer-ai/fla…
2
5
1
104
21,016
赠Zeyuan Allen-Zhu🫡 世言大模惟精调,君执初心问真章。 动力学中观涨落,表达性里辨行藏。 万里独征云和月,所慕岂在功与赏。 他年若问登临处,格物致理照新航。 突き破れ、扉の向こうへ! (source: 鋼の錬金術師 ED,《扉の向こうへ》)
I've decided it's time to resign from FAIR. 🫡 I'm especially grateful for the compute that made much of my research possible: 400 H100/200 GPUs allocated to me by FAIR; over a thousand H100s borrowed from FAIR Europe's CodeGen team led by Gabriel (@syhw ); thousands more borrowed from a legacy cluster in early 2025; and access to thousands of idle and low-prio GPUs across FAIR. For clarity, since joining FAIR four years ago, I've been on FAIR-level pay throughout. I mention this simply to avoid any confusion in any media coverage. Many thanks to @ylecun for founding FAIR, and to Joelle (@jpineau1 ) for steering FAIR through the later years of its golden age. じゃあね — so long, FAIR.
7
9
1
249
46,269
We are hiring engineers working on sandbox and LLM inference. If you are interested, please send your CV to [email protected]
24
34
5
668
53,980
🚨A novel way to do RL in LLM post-training! Inspired by our previous path-not-taken work (arxiv.org/abs/2511.08567), we dig deep into the learning trajectory of RL and find that optimizing singular vectors (i.e., rotation) of weight matrices suffices for good performance in RL. The resulting “isospectral optimization” reaches matched scores with substantially fewer training steps. Great work from @zhu_hanqing666 and the co-authors!
People keep asking me: what's different about optimization in RL? Seemingly nothing — the pre-training stack just works (Adam, even SGD 👀 @saagnikkk). Bringing some answers from my last work (sorry for the delay — been cooking 🚀). We introduce ISO: Isospectral Optimization: an RLVR-native optimization stack. Built on one simple observation, spectral inheritance: RLVR can reuse the base model's spectrum and acquire new behavior purely through the singular frames. 🧩 Offline: ISO-Merger — consolidates RL experts into one model with no data, no rollouts, no OPD. Checkpoints only. ⚙️ Online: ISO-Optimizer — a drop-in wrapper on AdamW / Muon that matches AdamW's accuracy with ~2.7× fewer steps on Qwen3-8B-Base. 📄 arxiv.org/abs/2607.19331 🌐 iso-rlvr.github.io/ 🧵👇
9
79
2
659
69,812
😂This has happened numerous times in the history. Human + smart phone = superhuman. Human + LLM = supersuperhuman. We are already way better than those miserable sapients 20 years ago.
Replying to @tydsh
The other dystopia: Everyone has access to open-weight superhuman brains inside superhuman robots and can tell them what to do.
4
32
10,418
Strong support for that. The worst situation that could happen to human kinds is that a few elites control the best model (and their APIs), while other people treat them as "god" and pray for access. That would be the real dystopia...
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models. images.nvidia.com/pdf/Open-W…
6
22
2
286
24,935
In AMD AI Conference today.
20
4,598
wow... looks correct to me?!
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
6
1
1
103
24,116
Great to have you @ChengleiSi in the team! Let's rock together😆!
Just want to add that @tydsh himself is a great example of down-to-earth do-er! Last week he asked me for some model checkpoints, and immediately ran a bunch of experiments to discover some important findings. You can’t be a great LLM researcher without being down-to-earth! 🫡
2
59
17,318