@BerkeleyML

Students at UC Berkeley working on academic research, ML education, industry projects, and fostering a vibrant ML community 🧠💡

Berkeley, CA
Joined December 2016
check out this paper from one of our alums!
Excited to share our work on a model for mapping molecules to olfactory receptors to perception, out today! Check out the link in thread ⬇️
7
693
check out this paper on block sparse attention from two of our members!
Hardware yearns for block sparse attention, yet it seems largely absent from open weight LLMs. DeepSeek developed NSA, and people speculated DeepSeek v4 would integrate it, yet it was never utilized. We have a hypothesis as to why. We found that replacing dense attention with NSA significantly degraded its ability on synthetic retrieval tasks, even when finetuned on them. On our 32k context benchmark, it scored 0.300 compared to dense attention’s 0.904. We found the reason, and how to fix most of it. Introducing COBS🌽(Cumulant Order Block Sparse Attention), with @AdiGhai18 @sanjitneelam @ZVasania @tensorpro: • Raises NSA’s 0.300 → 0.820, closing ~86% of the gap to dense • 15.15x less KV-cache read traffic than dense (just 1.21x the NSA baseline) • Lower position-wise NLL than dense in our comparison The key insight: block selection is the keystone to block sparse attention, and existing methods are mathematically stuck in storing a first-order approximation of a cumulant generating function. COBS caches a compressed second cumulant and escapes this ceiling. Details in the paper.
6
1,182
are you guilty of --dangerously-skip-permissions? check out this new intervention framework by our alum!
Deciding when to jump in and help someone—and when to hold back and let them work through it—is something humans navigate constantly. How do AI assistants handle this tradeoff? We introduce Int-Bench, a framework for evaluating interventions during problem-solving tasks.
8
1,116
Check out this Mixture of Experts project by our member!!
This week, I wanted to see if we could get the smallest possible Mixture of Experts going that takes a few hours to pretrain on a single GPU, is fast for experiments, but also somehow performs well compared to larger MoEs. For reference, we have Nanochat for transformer experimentation that's 561M params. But Qwen's smallest MoE is still 14 billion total params which is still +++ GPUs. It's been very fun to work on nanochat version of that 🐸⚡️ nanoMoE is a 500M total param MoE that you can train for <5 hours on a single H100 using 3.5-4 less OOMs than the smallest MoEs. And packs a punch for its size (reaches 87% accuracy of OLMoE-1B-7B.). Also tried out the new Quantile Balancing from kimi so no hyperparam sweeping. More numbers on performance here: github.com/chloechiaw/nanomo…
7
974
Check out this awesome explainer of DiffusionGemma by one of our members!
This is Google’s new diffusion LLM, DiffusionGemma’s denoising canvas over time. Diffusion LLMs can generate tokens in flexible order. But in practice, do they just become autoregressive anyway? 1/ 🧵
9
671
Machine Learning at Berkeley retweeted
What if the best visual reasoning steps are ones humans can’t specify? 🤔 Existing VLM reasoning is often constrained by language, pixels, and human-designed intermediates. We introduce Latent Implicit Visual Reasoning, where we show that VLMs can discover the best visual reasoning steps by themselves — no bboxes, no intermediate images, no extra supervision. Presenting this week at CVPR! (1/n)🧵
2
9
3
34
8,858
our members ran a great reading group on major architecture changes in deepseek v4!! they covered: hyper connections + manifold hyper connections, KV cache + MQA/GQA intro, Deepseek Sparse attention (prerequisite to understanding the new CSA) & a walk through of CSA and HCA
2
7
603
join us for the last bioml seminar of this semester!!
We'll be closing out this semester's Berkeley BioML Seminar on 4/28 with a talk from @antoinekoehl on PEINT, a powerful deep learning framework for both phylogenetic inference and protein engineering. Sign up below! luma.com/3fbbftwy
2
4
691
super proud of members Avy Harish (@AvyukthH60737), Sahir Tandon and Chris John for hosting our first physics informed ml seminar! they covered: physics-informed nns, neural odes, sparse identification of nonlinear dynamics, fourier neural operators & other applications
1
1
1
3
777
Machine Learning at Berkeley retweeted
We're excited to announce the fourth Berkeley BioML seminar of the semester happening next Tuesday 4/7! Join us for a talk by Shreshth (@shreshth_gandhi) from @tahoe_ai on Scaling Perturbation-Trained Single-Cell Foundation Models! luma.com/njhd2xga
5
9
1,643
> intuitions of how to go from pytorch to jax > training language models from scratch > how to scale ur model pt 2 > lfggg
Wrote a deep dive on implementing a language model from scratch in JAX and scaling it with distributed training! If you’re coming from PyTorch and want to see how the same ideas look in JAX, or just want a hands-on intro to distributed training, check out this blog post: chuyishang.com/blog/2026/jax… Comes with code + an assignment and test cases so you can follow along!
1
4
30
4,724
who are the best people to talk to about physics + ml?
1
1
4
891
Machine Learning at Berkeley retweeted
We're excited to announce the third Berkeley BioML seminar of the semester happening next Tuesday 3/2! Join us for a talk by Vivek Natarajan (@vivnat ) from GDM on advancing science and medicine with collaborative AI agents. luma.com/532xzjiu
6
11
1,849
Join us for a tech talk with Transluce on 2/23 from 6-7 pm, in Cory 540AB! Transluce is a non-profit AI lab working to ensure that AI oversight scales with AI capabilities. (Pizza provided!)
1
1
4
530
We’ll use Docent, our tool for agent oversight that analyzes execution traces at scale, to dig into what’s really happening. Folks should join for an interactive demo and a chance to dig in themselves.
1
136