@Azlock1729

Researcher @logic_int | PhD @GatsbyUCL

London
Joined December 2025
New paper w/ @SaxeLab & Nishil Patel! 🧵 Does RL post-training teach models anything new, or just amplify skills already in the base model? We built a fully auditable testbed to settle it — and caught RL composing new strategies in the act.
5
4
53
6,675
4/ Rejection fine-tuning? Improves early, then plateaus. The difference isn't exploration volume — it's selectivity. RFT churns out shortcut-y, often invalid rewrites. RL concentrates its exploration into valid, reusable structure.
1
3
511
3/ The result: with only a final-answer reward, RL solves held-out problems the base model basically never solves. And we can see how: it first sharpens primitive skills, then builds composed procedures out of them, and reuses them as a stable toolkit.
1
1
13
1,815
2/ Why this is hard to prove in LLMs: you never know what was in pretraining. Our fix — a rewrite-grammar world where the pretraining distribution is fully known and every rewrite the model makes can be verified.
4
397
1/5 Excited to share LaMo: A Latent Motion World Model for Long-Horizon Prediction, to be presented at the ICLR 2026 Workshop on World Models. LaMo predicts compact latent motion rather than the next dense latent state.
2
1
9
5,618
5/5 The broader takeaway: For world models, the right primitive may not be “predict the next observation.” It may be: learn a latent state, learn the motion that transforms it, and roll out by composing motions.
1
2
197