@rosinality

ML Engineer

London, United Kingdom
Joined October 2008
I post the papers I find interesting. There are so many papers published these days, and I frequently miss great papers. I appreciate paper recommendations via DM, but I tend to only post papers I discover on my own to keep my list personally curated.
6
3
1
117
17,277
You can drop the vision encoder if pretraining compute is larger than 1e22 flops.
4
34
3
523
29,482
More efficient setup for the looped MoE. 2x experts, 0.5x looped layers, 2x loops, and attention untying. Now the looped transformer becomes a problem closer to better allocation of resources instead of inductive biases.
2
13
101
4,942
Rosinality retweeted
Added "Auto-character Coverage" to SentencePiece as a clean alternative to Byte-Level BPE (BBPE). It globally optimizes the vocabulary budget without invalid UTF-8 byte fragments, achieving comparable or better compression. google.github.io/sentencepie…
1
11
42
4,796
They are still avoiding using synthesized data, but maybe they have used data from more capable models? But for multimodal data they don't use synthetic data (which is the area where synthetic data is used extensively). They now explicitly mention the scaling ladder, their own crawling system.
1
23
2,628
As cited in the technical report it is a decoder-decoder architecture for KV cache compression (arxiv.org/abs/2405.05254). Very interesting.
Deepseek V4.1 Flash 552B total, 8/16B active with a new arch trained on 45T tokens, there are different active parameters for input/output tokens with the encoder/decoder arch, engram, new sparse attention, new mHC, native vision very high benchmarks (beating K3), insane efficiency, and as always amazing tech report this is probably the most novel arch i've seen in a while, pretty insane
2
9
1
63
9,147
I can't understand why people keep trying to say some company won the race after each model release. What is important for the model company is whether they have a roadmap and good direction and whether they are able to achieve it, not the model at each specific time point which eventually gets deprecated soon, and these things are generally hard to know from the outside of the company.
1
1
34
1,958
Maybe full-bandwidth transformer (not looped transformer) style architecture could allow more obscure cot? Though I think it will still be anchored around discrete tokens.
5
27
2,291
Great results when everyone talks about looped transformers. Looped transformers are now more compute-efficient compared to non-looped ones.
Can Looped Transformers still help when FLOPs, parameters, and KV cache are all matched? We introduce SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers. The answer is yes. And the advantage grows with scale. 🧵 1/8 Paper: arxiv.org/abs/2609.01343
1
5
1
42
5,648
Rosinality retweeted
Can Looped Transformers still help when FLOPs, parameters, and KV cache are all matched? We introduce SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers. The answer is yes. And the advantage grows with scale. 🧵 1/8 Paper: arxiv.org/abs/2609.01343
8
71
7
451
74,700
arxiv.org/abs/2608.30627 Inserting reasoning tokens into the pretraining data. This has been tried multiple times, but how scalable is it?
3
9
68
5,482
arxiv.org/abs/2608.24814 Effective learning rate, the ratio of learning rate and weight norm, governs training dynamics. This could allow transferring the settings across norm control methods (by matching effective learning rate).
3
18
167
10,796
arxiv.org/abs/2608.20061 Hyperparameter transfer attempt for 10T scale. Transfer over token horizon was done through a scaling law.
1
7
109
6,142
arxiv.org/abs/2608.19197 Synthetic environment generation through solver agent and environment generator dynamics. The rewards for environment generator are calculated using the difference of solver rewards conditioned on privileged information or not.
2
9
101
5,688
arxiv.org/abs/2608.18486 Introducing cross-layer connections is popular now. The problem is how to parallelize it (like Jacobi iterations) and whether it is enough for large-scale training.
14
113
6,586
arxiv.org/abs/2608.17981 Recurrent transformer that injects activations from the top layers to the bottom layers at the next step (arxiv.org/abs/2608.08888). Why does this work without training?
32
105
32
1,058
235,934
arxiv.org/abs/2608.17286 Scaling law estimation for text-to-image diffusion. One interesting result is that diffusion is forgiving for overtraining in the sense that the loss difference between compute optimal and overtrained models is relatively small.
8
1
71
4,630
arxiv.org/abs/2608.14071 Data repetition during pretraining, when non-repeated data is available to fill the remaining portion to keep TPP constant. It is another observation on how high quality data could be repeated more, with the twist that a larger model (!) and shorter LR decay tolerate more repetition better. This could interact with the "non-repeated" web data part, as it could be a balance of noise fitting between noisy unique data and high quality repeated data.
5
24
1
204
20,237
arxiv.org/abs/2608.11669 Randomly drop rubric items to suppress reward hacking.
20
2
201
12,594