@ezyang

I work on PyTorch at Meta. Chatty alt at @difficultyang.

Edison, NJ
Joined May 2008
When counting FLOPs do you find SI units or scientific notation more intuitive?
35%SI units (2 EFLOPs)
65%Sci. notation (2 x 10^18)
234 votes • Final results
6
11
4,158
A small observation. Conventionally, SAC is applied at a transformer block granularity. However, many SAC policies end up always saving the result of a residual add. Since these are the sole inputs into the attn/mlp, it works to just SAC these individually.
1
5
1,777
Because you usually want to recompute the norms, it's a little inaccurate to say we SAC the "attn/mlp", really you want to SAC the norm + attn/mlp.
487
Understated benefit of AIs for code review is you can easily tell if the diff author actually handled your comments or not
1
12
1,940
Having an incredible amount of analysis paralysis on 03 of the DSv3 series
5
1,876
For reasons 👀 I had Astra cook up the roofline chapter of "How to Scale Your Model" as an interactive edition: ezyang.github.io/interactive… This edition has NVIDIA GPUs as well as split up HBM/Network analysis. Hope someone finds it useful.
4
23
376
18,127
A shower thought: it's never been easier to create one off dynamic analyses. Design the abstract domain, use an LLM to annotate all of the terminals in your program, and then run the analyses to check non-local problems.
1
2
1
14
5,779
In an unusual turn of events, I'll be in the Bay Area the weekend of Oct 17-18 (before PTC) without any plans or family obligations. Would be fun to meet up with mutuals while I'm around; DM me if you are interested!
9
1,457
It's like KOLMOGOROV COMPLEXITY but for LARGE LANGUAGE MODEL PROMPTS
1
28
3,144
I knew this, but some recent stuff I've been working on has really emphasized that type systems are not solely just about "finding bugs"
2
1
52
4,714
So it seems, if Codex doesn't show in mobile, the best way to get to your remote control session is to go to your profile, go to remote control, wait for it to connect to the active session, and then click the notification to get to the Codex UI
3
1
19
3,677
X articles is where text goes to die, since it doesn't get indexed anywhere else
1
47
2,826
So the answer is, apparently, that the feature hasn't been officially released, which is why I can't find it in the main app menu lmao
One of my favorite features of Codex/ChatGPT is not even documented in the main docs, I discovered it on Reddit. So basically we have a GPU slurm cluster that we SSH into, and I wanted to be able to run Codex on it from my phone. ChatGPT Desktop supports SSH connections but how about accessing it directly from mobile app? Turns out there is a separate "Remote Control" feature not officially documented anywhere. Simply run "codex remote-control start" on the server you want to Codex into from your mobile. Then run "codex remote-control pair" to getting pairing codes. Then in the app, go to the Remote section, "Add connection", select "Pair manually", and put in the pairing code. It's that simple! I use this feature pretty much every day to manage autoresearch experiments and other training runs on our GPU cluster.
9
1
1
92
27,452
One of the weirder bugs: the ChatGPT app on Android never shows me the Codex menu entry, even though it will drop me in there if I manually pair a remote session
1
13
2,283
There's definitely some sort of comparative advantage thing going on between me and LLMs on generating text, because no matter how good the LLMs get at generating text I just generate the text they're not good at
1
1
32
3,339
This is now fixed and errata'd in the H100 memory study, thank you Vlad for reporting.
great stuff!! one nit on the expert diagram: I think swiglu doesn't consume e4m3 directly since DeepGEMM outputs bf16, and they probably fuse quantization into the swiglu kernel deepseek open sourced their perfetto trace for v3 in early 2025 github.com/deepseek-ai/profi…. Lots of cool stuff there, for example, the swiglu time matches topK*seqlen*intermediate*(2 + 2 + 1 + 1 + 1)/2.8e12 = 8*4096*2048*7/2.8e12 = 0.16ms. IMO most likely they read two bf16 tensors and output three fp8 tensors: two of them are probably quantized inputs, and the third one is quantized `silu(a)*b` to feed into the next DeepGEMM call. (2.8e12 above is roughly 85% from 3.35 TB/s from the spec on hoppers)
37
4,783
Announcing "DeepSeek-V3: from roofline to reality", a blog post series that digs into roofline analysis specifically on DSv3. We're kicking off this series with two posts! deepseek-v3.ezyang.com/index…
5
47
4
497
33,347
This post series would not at all have been possible without Fable, who was my tireless data visualization engineer and built all of the visualizations and animations. However, except for small exceptions (like some of the tooltips/labels), all of the text is human written.
1
1
7
768
P.S. There are existing sims that take your model+parallelism and tell you what kind of MFU you should expect. While we will eventually end up building one of these, I care more about the journey (How is the sim put together?) than the destination. Intuition over tools!
1
7
684