@rosinalityi
iAccount based inUnited Kingdom
About this account
- Account based in
- United Kingdom
- Connected via
- Korea App Store
Account-level information from X, not a live location or the device used for a specific post.
ML Engineer
London, United Kingdom
Joined October 2008
- Tweets33.3K
- Following1K
- Followers8K
- Likes23.5K
Pinned Tweet
I post the papers I find interesting. There are so many papers published these days, and I frequently miss great papers. I appreciate paper recommendations via DM, but I tend to only post papers I discover on my own to keep my list personally curated.
More efficient setup for the looped MoE. 2x experts, 0.5x looped layers, 2x loops, and attention untying. Now the looped transformer becomes a problem closer to better allocation of resources instead of inductive biases.
Added "Auto-character Coverage" to SentencePiece as a clean alternative to Byte-Level BPE (BBPE).
It globally optimizes the vocabulary budget without invalid UTF-8 byte fragments, achieving comparable or better compression. google.github.io/sentencepie…
They are still avoiding using synthesized data, but maybe they have used data from more capable models? But for multimodal data they don't use synthetic data (which is the area where synthetic data is used extensively). They now explicitly mention the scaling ladder, their own crawling system.
As cited in the technical report it is a decoder-decoder architecture for KV cache compression (arxiv.org/abs/2405.05254). Very interesting.
As cited in the technical report it is a decoder-decoder architecture for KV cache compression (arxiv.org/abs/2405.05254). Very interesting.
Deepseek V4.1 Flash 552B total, 8/16B active with a new arch trained on 45T tokens, there are different active parameters for input/output tokens with the encoder/decoder arch, engram, new sparse attention, new mHC, native vision
very high benchmarks (beating K3), insane efficiency, and as always amazing tech report
this is probably the most novel arch i've seen in a while, pretty insane
I can't understand why people keep trying to say some company won the race after each model release. What is important for the model company is whether they have a roadmap and good direction and whether they are able to achieve it, not the model at each specific time point which eventually gets deprecated soon, and these things are generally hard to know from the outside of the company.
Maybe full-bandwidth transformer (not looped transformer) style architecture could allow more obscure cot? Though I think it will still be anchored around discrete tokens.
Great results when everyone talks about looped transformers. Looped transformers are now more compute-efficient compared to non-looped ones.
Can Looped Transformers still help when FLOPs, parameters, and KV cache are all matched?
We introduce SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers.
The answer is yes. And the advantage grows with scale. 🧵 1/8
Paper: arxiv.org/abs/2609.01343
Rosinality retweeted
Can Looped Transformers still help when FLOPs, parameters, and KV cache are all matched?
We introduce SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers.
The answer is yes. And the advantage grows with scale. 🧵 1/8
Paper: arxiv.org/abs/2609.01343
arxiv.org/abs/2608.30627
Inserting reasoning tokens into the pretraining data. This has been tried multiple times, but how scalable is it?
arxiv.org/abs/2608.24814
Effective learning rate, the ratio of learning rate and weight norm, governs training dynamics. This could allow transferring the settings across norm control methods (by matching effective learning rate).
arxiv.org/abs/2608.20061
Hyperparameter transfer attempt for 10T scale. Transfer over token horizon was done through a scaling law.
arxiv.org/abs/2608.19197
Synthetic environment generation through solver agent and environment generator dynamics. The rewards for environment generator are calculated using the difference of solver rewards conditioned on privileged information or not.
arxiv.org/abs/2608.18486
Introducing cross-layer connections is popular now. The problem is how to parallelize it (like Jacobi iterations) and whether it is enough for large-scale training.
arxiv.org/abs/2608.17981
Recurrent transformer that injects activations from the top layers to the bottom layers at the next step (arxiv.org/abs/2608.08888). Why does this work without training?
arxiv.org/abs/2608.17286
Scaling law estimation for text-to-image diffusion. One interesting result is that diffusion is forgiving for overtraining in the sense that the loss difference between compute optimal and overtrained models is relatively small.
arxiv.org/abs/2608.14071
Data repetition during pretraining, when non-repeated data is available to fill the remaining portion to keep TPP constant. It is another observation on how high quality data could be repeated more, with the twist that a larger model (!) and shorter LR decay tolerate more repetition better. This could interact with the "non-repeated" web data part, as it could be a balance of noise fitting between noisy unique data and high quality repeated data.
Interesting.
github.com/cordiverse/paper
DEEPSEEK HARNESS IS OUT
github.com/deepseek-ai/deeps…
Maybe this is originally used for this? koishi.chat/en-US/ Touhou is everywhere.
arxiv.org/abs/2608.11669
Randomly drop rubric items to suppress reward hacking.