@jonathanrichens

Research scientist in AI safety @GoogleDeepMind

Joined August 2020
Are world models necessary to achieve human-level agents, or is there a model-free short-cut? Our new #ICML2025 paper tackles this question from first principles, and finds a surprising answer, agents _are_ world models… 🧵
35
173
29
1,063
186,615
I'm incredibly excited to announce the founding of the Mathematical AI safety Institute (MAISI) maisi.org/. AI safety needs more foundational theoretical development, and mathematicians have the skills and the mindset to help! MAISI is an independent institute with visiting positions ranging from 1 semester to 2 years. Our goal is to get mathematicians up to speed and working on research directions in AI safety as quickly as possible. There is important work to be done, and there is real progress to be made. YOU can help! Applications are open now! MAISI is aiming to hire 10-30 mathematicians to join us in the Bay Area by January, and scale up to 30-100 for September 2027. If you're a mathematician interested in channeling your skills toward the most important problem of our time, please apply today! maisi.org/apply
38
213
36
1,184
210,295
Applications for my SPAR project are open (til next Tues). Come work with me on weird agent foundations / causality stuff for a couple months 🤖. No coding, just scribbling. sparai.org/projects/f26/recC…
We're excited to announce that applications are now open for the Fall round of SPAR! 🧵 Join us for 3 months (September - December) in a remote program to work on impactful research projects with expert mentors in AI safety, AI policy, AI security, and Biosecurity.
1
3
17
1,389
Predicting the answer to interventional "what if?" questions — the outcome of an action you never took — need a *mechanistic* model, not a curve fit. And you can only learn one by *experimenting*. Experiments are costly, so the real game is **data efficiency**. Meet the Model Discovery Agent (MDA). 🧵
29
173
28
1,199
207,935
100 years later, computer science has caught up with Joyce
This quoted post is unavailable.
1
10
1,332
Jon Richens retweeted
🫡 After 7 years, I've just left Google Deepmind to start an AI Safety nonprofit, around Scalable and Human Oversight (i.e. building stronger "judges")! 🇨🇦 And, I'll be at FAccT in Montreal this week! (1/4🧵)
96
96
18
2,142
178,166
Turns out you can invert the Bellman equation to recover an agent's world model from its value function. Excited by the potential applications of this work, lead by @_aletcher. My fave bit - RL agents implicitly model latent variables they were never trained to optimize for..🧵
Model-free agents learn to maximise reward without modelling the environment. Right? In recent work, we challenge this narrative by proving that agents, trained on a sufficiently rich set of goals, encode a unique and accurate world model in their value functions. 1/
9
63
4
496
74,860
Jon Richens retweeted
I'll very soon be hiring a postdoc at UCL for a 2y project combining interpretability methods with behavioural evals to study whether goal and belief representations can be reliably extracted and manipulated in LM agents. More details soon. Please get in touch if interested.
4
28
2,963
Theory + alignment ❤️. Really excited to see what comes out of this new org.
We are starting a new, nonprofit alignment organization, ⊢ Sequent Research, bringing together researchers previously on UK AISI’s Alignment Team, Timaeus, and elsewhere to research how to align superintelligence. We are hiring! 🧵
13
1,115
Jon Richens retweeted
🚨Transformers don't learn Newton's laws? They learn Kepler's laws! Like us, transformers don't predict a flying ball via a differential equation, but by fitting a curve. Moreover, reducing context length steers a transformer from Keplerian to Newtonian. Compression in play.
24
203
27
1,187
118,361
Jon Richens retweeted
Thrilled to share our new #NeurIPS2025 paper done at @GoogleDeepMind, Plasticity as the Mirror of Empowerment We prove every agent faces a trade-off between its capacity to adapt (plasticity) and its capacity to steer (empowerment) Paper: david-abel.github.io/plastic… 🧵🧵🧵👇
25
68
12
447
103,357
I will be a SPAR mentor this Fall🤖 Check out the programme and apply by 20 August to work with me on formalising and/or measuring and/or intervening on goal-directed behaviour in AI agents More info on potential projects here 🧵
1
3
2
12
2,478
2 years ago, @ilyasut made a bold prediction that large neural networks are learning world models through text. Recently, a new paper by @GoogleDeepMind provided a compelling insight to this idea. They found that if an AI agent can tackle complex, long-horizon tasks, it must have learned an internal world model—and we can even extract it just by observing the agent's behavior. I wrote a blog post unpacking this groundbreaking paper and what it means for the future of AGI 👇 richardcsuwandi.github.io/bl…
20
106
14
979
88,636
Can we trust a black-box system, when all we know is its past behaviour? 🤖🤔 In a new #ICML2025 paper we derive fundamental bounds on the predictability of black-box agents. This is a critical question for #AgentSafety. 🧵
4
23
1
122
36,140
Causality. In previous work we showed a causal world model is needed for robustness. It turns out you don’t need as much causal knowledge of the environment for task generalization. There is a causal hierarchy, but for agency and agent capabilities, rather than inference!
3
2
39
4,057
… and many more! Check out our paper arxiv.org/pdf/2506.01622, or come chat to me at #ICML2025. Joint work @GoogleDeepMind with @dabelcs, @alexis_bellot_, @tom4everitt
5
4
2
43
3,687