@GPU_MODE

Your favorite GPU community

gpumode.com
Joined September 2024
A very interesting blog post by @EliaIsPosting about applying autoresearch to a @GPU_MODE competition, winning 3rd place! Setup includes an orchestrator + subagents: idea generation, workers to implement the ideas, redteamers to audit the conclusions. Measuring actual signal and not regular variance is obviously a major challenge, which is highlighted in this blog post. The solution was an escalating evaluation protocol, which brings the comparison floor to ~0.15%. Link to blog post: elianaive.com/posts/another-…
7
6
1
40
7,382
GPU MODE retweeted
oh shit i am number 1 on hackernews atm (3x hackernews now lfg)
my latest blog post "auto-research with codex: how I achieved a 212x faster kernel over baseline with codex in GPU Mode's qr_v2 problem" is up now. in this post, i talk about my approach towards auto-kerneling on the QR decomposition problem. sankalp.bearblog.dev/autores…
25
34
2
1,043
108,475
Today, our CTO @ChrisKitching17 joins @GPU_MODE - one of the sharper corners of the accelerated HPC community - for a live talk on YouTube. 📅 7-8:30PM London / 11AM-12:30PM SF Streamed via GPU MODE's Discord. If you spend your day in kernels, compilers, or wrestling GPUs from more than one vendor, come watch.
1
2
8
589
GPU MODE retweeted
(1/9) I'm thrilled to share the open-source release of Mixture-of-Kittens (MoK), our MoE megakernel for NVL72s! MoK fuses all mixture-of-experts communication and computation into a single, fully deterministic kernel, and powers Composer training across tens of thousands of GPUs. Joint work with @nash_c_brown, @hmwildermuth, @tmwilliamlin168, and @ellev3n11
We're open-sourcing Mixture-of-Kittens (MoK), our MoE training megakernel for NVL72s. It fuses all Mixture-of-Experts communication and computation into a single, fully deterministic kernel, and runs up to 2.37x faster than the strongest public baselines.
18
45
4
438
46,996
Kinda wild that our community got together and collectively made Kimi inference on MI355X faster than B200 across all settings. Congrats to the winners team RadeonFlow and thanks again to AMD for working so closely with our community
3
8
3
82
11,872
The most important systems projects are all open
I stand on the shoulders of giants who stood on the shoulders of giants. So much of the progress we’ve made has been because of the openness of our great predecessors. Proud to support microsoft.com/en-us/corporat…
17
3,668
If you'd like to know more, here's a nice talk at @GPU_MODE introducing the One Layer Deeper competition.🙂 youtube.com/watch?v=ustDJbxO…
Played around with this a bit. It’s nice to see folks thinking about it, and it is certainly an important problem to address. It's a deceptively good benchmark that asks whether gradient-based learning can learn a procedure for reusable computation that generalises to more steps than it saw during training, or whether it just learns a map between inputs and outputs. The question is whether the model can learn the serial computation itself, rather than relying primarily on test-time compute (in token space), and whether latent depth can provide a cheaper form of sequential computation. You can see the thought that has gone into designing the tasks. Kudos to folks @coreauto and @tilderesearch.
2
4
1,661
We'll be doing a Q&A and demo tomorrow at 3:00pm Tuesday July 21 on our youtube channel. Please bring your questions!
Can a model run deeper at test time than it was ever trained to? And if depth becomes a loop instead of a stack, do we need better optimizers to keep training stable? @coreauto is collaborating with @tilderesearch on an optimizer x architecture competition: "One Layer Deeper" We picked a cursed problem y = x^(2^T) mod N where each squaring depends on the last. A vanilla transformer has a fixed depth so it falls flat on its face the moment T exceeds the serial compute it can do in one forward pass. A fun reminder that architectures have inductive biases. This isn't quite a nano-gpt speedrun experience and it's not a typical @GPU_MODE kernel competition either, it's closer in spirit to the LLM efficiency competition I worked on many years ago with @weiwei_msr and @Jisaacso in that it's open ended in a specific way. The UX is kernelbot like, your submission is a single python file that defines your model, optimizer and loss. You don't need a GPU to participate at all since we're reusing a queue based job system powered by @modal and @northflank There's tons of unexplored ideas and I'm really not sure what will win out but I'd be particularly excited to see people playing with new optimizers, adaptive depth and depth extrapolation. Since this is a new format, we'll work with the community to refine the rules for another week before we start for real. Submissions are open now!!
8
2,941
Great results thus far from the 9 participants in the @GPU_MODE competition. 11B tokens and $500 in @modal credits thus far. Thanks to @damngamerz and everyone else for participating!
I'll be leading a sprint on the @GPU_MODE Cholesky competition at @europython / @EuroSciPy this Sat/Sun in Krakow! Come learn how to autoresearch and write CUDA kernels. @modal has generously donated GPU access and @nvidia will provide frontier LLM access for the sprint.
6
6
50
7,457
Can a model run deeper at test time than it was ever trained to? And if depth becomes a loop instead of a stack, do we need better optimizers to keep training stable? @coreauto is collaborating with @tilderesearch on an optimizer x architecture competition: "One Layer Deeper" We picked a cursed problem y = x^(2^T) mod N where each squaring depends on the last. A vanilla transformer has a fixed depth so it falls flat on its face the moment T exceeds the serial compute it can do in one forward pass. A fun reminder that architectures have inductive biases. This isn't quite a nano-gpt speedrun experience and it's not a typical @GPU_MODE kernel competition either, it's closer in spirit to the LLM efficiency competition I worked on many years ago with @weiwei_msr and @Jisaacso in that it's open ended in a specific way. The UX is kernelbot like, your submission is a single python file that defines your model, optimizer and loss. You don't need a GPU to participate at all since we're reusing a queue based job system powered by @modal and @northflank There's tons of unexplored ideas and I'm really not sure what will win out but I'd be particularly excited to see people playing with new optimizers, adaptive depth and depth extrapolation. Since this is a new format, we'll work with the community to refine the rules for another week before we start for real. Submissions are open now!!
14
58
10
497
77,895
Second problem is now out: dense symmetric eigenproblem A=QΛQT. Solution due on July 15! We've also enabled ncu profiling for your agents on a @verdacloud cloud box sponsored by our good friends at Brev at @NVIDIAAI
Launching a new kernel competition: Linear Algebra Kernels For The Age Of Research. First problem: batched QR decomposition on B200. Old math, modern hardware. Prize: Rare swag and hangout in SF
6
15
4
156
85,949
GPU MODE retweeted
We taught a brand-new mini-series this year at @SCSatCMU on Modern GPU Programming for ML Systems, as part of the ML Systems course, touching on fun questions like what data layout swizzling is, how to use 3D TMA, and state-of-the-art Blackwell programming. We released a curated online book based on the materials: mlc.ai/modern-gpu-programmin… check it out
24
291
6
1,824
204,480
This isn't one problem - it's a dozen or more! Specialized kernels for different shapes & kinds of matrices will win the day. Leveraging tensorcores & a mixture of numeric formats (FP64, FP32, TF32, FP16, FP8) is likely key to top answers. This may even end up memory bound.
We've released the QR problem, a more robust qr_v2 with a fresh leaderboard so please resubmit! Thank you to @blelbach, @myainotez and @nikhilbarhate99 for sharing feedback. Sorry if I missed anyone! I considered automatically backfilling all submissions but the rankings do change quite a bit so I figured a refresh would be better. Changelog * Fail submissions if they fail when we change random seeds * Add nasty correctness cases with more degenerate inputs in mixed batches * Recheck correctness when doing perf testing to avoid Volkswagen cheat * Reject Nan/Inf residuals * Validate each matrix factorization residual, since averaging was hiding bad matrices * Old QR is still open so folks can't see submissions but you can't submit anything to it Wontfix * Stream hacking is still banned via very blunt ban of the word "stream" we don't have a good solution for this * CUDA graphs are allowed but not particularly interesting to us Best submissions so far if I resubmit their solutions are
4
4
1
82
19,088
GPU MODE retweeted
one can dream
7
17
4
329
29,552
RoboBryce reigns supreme. There's some crazy beautiful stuff in there.
We've released the QR problem, a more robust qr_v2 with a fresh leaderboard so please resubmit! Thank you to @blelbach, @myainotez and @nikhilbarhate99 for sharing feedback. Sorry if I missed anyone! I considered automatically backfilling all submissions but the rankings do change quite a bit so I figured a refresh would be better. Changelog * Fail submissions if they fail when we change random seeds * Add nasty correctness cases with more degenerate inputs in mixed batches * Recheck correctness when doing perf testing to avoid Volkswagen cheat * Reject Nan/Inf residuals * Validate each matrix factorization residual, since averaging was hiding bad matrices * Old QR is still open so folks can't see submissions but you can't submit anything to it Wontfix * Stream hacking is still banned via very blunt ban of the word "stream" we don't have a good solution for this * CUDA graphs are allowed but not particularly interesting to us Best submissions so far if I resubmit their solutions are
1
1
1
36
5,271
GPU MODE retweeted
Gpu mode hackathon leaderboards always something like: - principal engineer at nvidia - dude named samhandwich_69 - and @myainotez
6
6
284
24,635
Launching a new kernel competition: Linear Algebra Kernels For The Age Of Research. First problem: batched QR decomposition on B200. Old math, modern hardware. Prize: Rare swag and hangout in SF
I have some mixed feelings about this result: On the one hand, it's genuinely impressive. I didn't know that Shampoo could be configured to perform this well on the benchmark. On the other hand, the way this performance boost was achieved seems difficult to call "Vanilla," for the following reason: According to @_arohan_, the boost depends upon fixing a numerical linear algebra issue that he observed to occur in my initial standard DistributedShampoo run. He fixed the issue by enabling the flag rank_deficient_stability_config=PseudoInverseConfig(). Here's the problem: This is an undocumented flag. It is contained within the 12,000-line DistributedShampoo codebase, but it does not appear in any user-facing documentation. As a result, if someone tries to train a model using DistributedShampoo without either (a) knowing about this special undocumented flag or (b) being prepared to detect and fix the numerical linear algebra issues that may occur without it, then they won't be able to achieve @_arohan_'s level of Shampoo performance. This level of effort would be considered atypical for mere hyperparameter tuning. -- [Note on Muon baseline in plot below: Rohan's post compared Shampoo to a slightly undertuned Muon baseline from 2026/05/01, which reached the target loss in 3375 steps. This resulted in a 50-step gap between Shampoo and Muon. In the figure below I'm using the up-to-date 2026/05/03 baseline, which reaches the target in 3325 steps. This results in the step-counts exactly matching between Muon and the tuned/stabilized Shampoo variant.]
12
32
15
412
169,334
GPU MODE has powered much of the public GPU kernel work online, with a permissive license from day one and generous credit from researchers, NVIDIA, AMD, and others. Today we’re moving our datasets to the Researcher Reciprocity License.
June 9th Researcher Reciprocity License "if you train on it, you let us generate - reverse terms of use void" Status quo 1. We teach frontier devs with ICLR/NeurIPS papers, OSS Github contributions 2. They use it to make frontier models 3. Then ban us from exploring our ideas We need a new license, original thinkers can't be an underclass to a tyrannical researcher fiefdom
8
26
3
421
34,113
GPU MODE retweeted
YEEEEEES
2
2
1
44
3,615