@BerkeleyMLi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Students at UC Berkeley working on academic research, ML education, industry projects, and fostering a vibrant ML community 🧠💡
Berkeley, CA
Joined December 2016
- Tweets250
- Following113
- Followers3.7K
- Likes289
check out this paper on block sparse attention from two of our members!
Hardware yearns for block sparse attention, yet it seems largely absent from open weight LLMs. DeepSeek developed NSA, and people speculated DeepSeek v4 would integrate it, yet it was never utilized.
We have a hypothesis as to why.
We found that replacing dense attention with NSA significantly degraded its ability on synthetic retrieval tasks, even when finetuned on them. On our 32k context benchmark, it scored 0.300 compared to dense attention’s 0.904.
We found the reason, and how to fix most of it. Introducing COBS🌽(Cumulant Order Block Sparse Attention), with @AdiGhai18 @sanjitneelam @ZVasania @tensorpro:
• Raises NSA’s 0.300 → 0.820, closing ~86% of the gap to dense
• 15.15x less KV-cache read traffic than dense (just 1.21x the NSA baseline)
• Lower position-wise NLL than dense in our comparison
The key insight: block selection is the keystone to block sparse attention, and existing methods are mathematically stuck in storing a first-order approximation of a cumulant generating function. COBS caches a compressed second cumulant and escapes this ceiling.
Details in the paper.
are you guilty of --dangerously-skip-permissions? check out this new intervention framework by our alum!
Check out this Mixture of Experts project by our member!!
This week, I wanted to see if we could get the smallest possible Mixture of Experts going that takes a few hours to pretrain on a single GPU, is fast for experiments, but also somehow performs well compared to larger MoEs. For reference, we have Nanochat for transformer experimentation that's 561M params. But Qwen's smallest MoE is still 14 billion total params which is still +++ GPUs. It's been very fun to work on nanochat version of that 🐸⚡️
nanoMoE is a 500M total param MoE that you can train for <5 hours on a single H100 using 3.5-4 less OOMs than the smallest MoEs. And packs a punch for its size (reaches 87% accuracy of OLMoE-1B-7B.). Also tried out the new Quantile Balancing from kimi so no hyperparam sweeping.
More numbers on performance here:
github.com/chloechiaw/nanomo…
Machine Learning at Berkeley retweeted
What if the best visual reasoning steps are ones humans can’t specify? 🤔
Existing VLM reasoning is often constrained by language, pixels, and human-designed intermediates.
We introduce Latent Implicit Visual Reasoning, where we show that VLMs can discover the best visual reasoning steps by themselves — no bboxes, no intermediate images, no extra supervision.
Presenting this week at CVPR!
(1/n)🧵
our members ran a great reading group on major architecture changes in deepseek v4!!
they covered: hyper connections + manifold hyper connections, KV cache + MQA/GQA intro, Deepseek Sparse attention (prerequisite to understanding the new CSA) & a walk through of CSA and HCA
join us for the last bioml seminar of this semester!!
We'll be closing out this semester's Berkeley BioML Seminar on 4/28 with a talk from @antoinekoehl on PEINT, a powerful deep learning framework for both phylogenetic inference and protein engineering. Sign up below! luma.com/3fbbftwy
super proud of members Avy Harish (@AvyukthH60737), Sahir Tandon and Chris John for hosting our first physics informed ml seminar!
they covered: physics-informed nns, neural odes, sparse identification of nonlinear dynamics, fourier neural operators & other applications
Machine Learning at Berkeley retweeted
We're excited to announce the fourth Berkeley BioML seminar of the semester happening next Tuesday 4/7! Join us for a talk by Shreshth (@shreshth_gandhi) from @tahoe_ai on Scaling Perturbation-Trained Single-Cell Foundation Models!
luma.com/njhd2xga
> intuitions of how to go from pytorch to jax
> training language models from scratch
> how to scale ur model pt 2
> lfggg
Wrote a deep dive on implementing a language model from scratch in JAX and scaling it with distributed training!
If you’re coming from PyTorch and want to see how the same ideas look in JAX, or just want a hands-on intro to distributed training, check out this blog post: chuyishang.com/blog/2026/jax…
Comes with code + an assignment and test cases so you can follow along!
Machine Learning at Berkeley retweeted
We're excited to announce the third Berkeley BioML seminar of the semester happening next Tuesday 3/2! Join us for a talk by Vivek Natarajan (@vivnat ) from GDM on advancing science and medicine with collaborative AI agents.
luma.com/532xzjiu
Join us for a tech talk with Transluce on 2/23 from 6-7 pm, in Cory 540AB! Transluce is a non-profit AI lab working to ensure that AI oversight scales with AI capabilities. (Pizza provided!)