@jvmncsi
iAccount based inUnited States!
About this account
- Account based in
- United States
- Connected via
- United States App Store
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
research @modal - ❤️s/RTs are randomized and differentially private
Brooklyn
Joined February 2010
- Tweets2.8K
- Following1.4K
- Followers1.2K
- Likes11.3K
we’ve been training some big models at decagon 👀
up to a trillion parameters on modal clusters. congrats on the launch!
Replying to @modal
Modal Clusters are now available.
Globally available, RDMA-connected nodes available instantly behind a single decorator.
modal.com/blog/modal-cluster…
jason retweeted
I’m so pumped to see Clusters make it into the wild.
2 years ago, we set out with an extremely ambitious goal: RDMA clusters, billed by the second, available instantly.
I didn’t always believe it was possible. But that didn’t stop us from trying :)
Replying to @modal
Modal Clusters are now available.
Globally available, RDMA-connected nodes available instantly behind a single decorator.
modal.com/blog/modal-cluster…
jason retweeted
so excited to see which reservations I’ll be able to get with my museinstinctdotgrokbot after I give it access to my calendar contact book email identity and also a gun
jason retweeted
a common problem with RL is reward hacking.
you could try to curb it with hard-coded rules or a non-deterministic judge.
or, make life easy and use Jev to make Qwen3.5-4B speak in alliterations that are actually good (at least i think so).
made with @modal and @typesafeai.
🫡
(pls merge my open PRs @sgl_project)
the unreasonable effectiveness of random search strikes again
“Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning”
KV cache eviction methods spend compute scoring which past tokens are important enough to keep.
Random Attention shows that for long reasoning, you can just preserve the prompt and randomly evict the rest independently across heads.
Across 4 models and 6 reasoning tasks, it matches the strongest prior methods while serving 32-43% more tokens/sec than TriAttention, suggesting reasoning traces are redundant enough that sophisticated token ranking often buys very little.
alphaxiv.org/abs/2609.03430
alt title: Stitch your rollouts together across clouds and regions
hey folks,
we are back with anotha london systems club. might be our best one yet.
@kennethnym @PrimeIntellect on how they built Prime Agent (~96% ARC-AGI 3 🤯)
Hamzah Chariwala @CallosumAI on benchmarking accelerators across different workload shapes
@jvmncs @modal on Stitch, our versioned control plane for disaggregated RL.
luma.com/0uhqgt6p
we ❤️ miles
Today we're launching Miles v0.1, an open-source RL framework for LLMs and multimodal models.
RL training is easy to start and hard to debug. Miles helps you ensure your run is correct, use hardware efficiently, and keep RL running at scale.
Over the past 9 months, 72 contributors have landed 1,326 commits, 85 GPU E2E CI tests, battle-testing Miles on frontier open models like Kimi K3, DeepSeek V4, Qwen 3.8, GLM 5.2, Inkling, MiniMax H3, etc.
Miles powers frontier-model development and production RL workloads at @humansand, @periodiclabs, @modal, @DecagonAI, @Eigent_AI, @nebiusai, @IBM and more, on both @NVIDIAAI and @AIatAMD hardware.
Here is what we built, and why teams picked Miles🧵
every day we stray further from the light
I've started writing my model code as a single, compilable `forward_backward` function:
1 - No more autograd. The backward math is written out, same as forward.
2 - No torch.nn module abstractions, just matrices and math. (Glorious!)
3 - Attach a `.gacc` to each parameter to store its gradients.
4 - Intermediate activations are explicitly stashed for use in backward.
The screenshot shows the activation stash for the 24-layer nanochat architecture, where Claude cleverly skipped a few with high footprint and low compute cost.
Most notably, the MLP needs the activation values both pre and post-ReLU, and they're 18GB each! We just store pre, and re-apply ReLU in backward.
(A recent modded-nanogpt PR did something similar-but-different, need to look at it more)
I'm not familiar with the backward math (Claude wrote it all), but I'm excited to gradually build up the same familiarity with it as the forward path.
“(somewhat) efficient, goal-driven retrievers of high-value strategies for gathering relevant context at inference time”
this type of “continual learning” is therefore more about learning to anticipate and react appropriately to non stationary tool responses without losing the general strategies that keep it all afloat
HOWEVER there is still a lot of work to make this function with LLMs. retraining weights or lora or w/e with RL while only changing data/reward
caps out quickly, and that’s prob why it’s a only really a usable system for handling small/steady amounts of distribution shift
btw this has been somewhat well known since at least 2017, when google released the TFX stack and publicized distillation alongside ~production-grade quantization and pruning in tensorflow serving
still so much to arbitrage from pre-LLM days for those who were around or care to study up
Continual learning is a bet that the retraining loop will get cheaper over time. With larger models, you can maybe run this loop once every few weeks. But with smaller models, you can run it nightly, per customer. And it keeps recursing: a model per company, then a model per client that company serves, then per matter. We’re getting closer to intelligence cheap enough to meter.
On the path to this, we received early access to, and post-trained @nvidia's Nemotron 3.5 Lightning on @harvey LAB. One click on the Trajectory platform, no new engineering. 0% to 8.3%, above Opus 4.6 at 6.6%.
jason retweeted
Ever since RL became a core part of LLM training, train–rollout numerical mismatch has been one of the recurring problems in RL infra.
There have been several great open-source efforts toward alignment, but production-scale support still often comes with tradeoffs in model scale, performance, or feature coverage.
Today, we're open-sourcing the deterministic train–rollout alignment stack we use for GLM-5.2-scale training — covering FP8 weights + FP8 KV rollout, DeepEP, DeepGEMM, and sparse attention.
github.com/THUDM/slime/pull/…
1/3
jason retweeted
We’re hiring ML research interns!
If you’re a current PhD student interested in problems like adaptive speculation, off-policy RL, elastic rollout fleets and more, come talk to us: jobs.ashbyhq.com/modal/38888…
Frontier models require frontier speculators. Our team moved mountains to achieve 460 tok/s with our custom DFlash model.
Try it day 0 at modal.com/endpoints.