@jonathanrichensi
iAccount based inUnited Kingdom!
About this account
- Account based in
- United Kingdom
- Connected via
- Web
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
Research scientist in AI safety @GoogleDeepMind
Joined August 2020
- Tweets153
- Following373
- Followers1.2K
- Likes364
Pinned Tweet
Are world models necessary to achieve human-level agents, or is there a model-free short-cut?
Our new #ICML2025 paper tackles this question from first principles, and finds a surprising answer, agents _are_ world models… 🧵
Jon Richens retweeted
I'm incredibly excited to announce the founding of the Mathematical AI safety Institute (MAISI) maisi.org/. AI safety needs more foundational theoretical development, and mathematicians have the skills and the mindset to help! MAISI is an independent institute with visiting positions ranging from 1 semester to 2 years. Our goal is to get mathematicians up to speed and working on research directions in AI safety as quickly as possible. There is important work to be done, and there is real progress to be made. YOU can help!
Applications are open now! MAISI is aiming to hire 10-30 mathematicians to join us in the Bay Area by January, and scale up to 30-100 for September 2027. If you're a mathematician interested in channeling your skills toward the most important problem of our time, please apply today! maisi.org/apply
Applications for my SPAR project are open (til next Tues). Come work with me on weird agent foundations / causality stuff for a couple months 🤖. No coding, just scribbling. sparai.org/projects/f26/recC…
Jon Richens retweeted
Predicting the answer to interventional "what if?" questions — the outcome of an action you never took — need a *mechanistic* model, not a curve fit. And you can only learn one by *experimenting*. Experiments are costly, so the real game is **data efficiency**.
Meet the Model Discovery Agent (MDA). 🧵
Jon Richens retweeted
🫡 After 7 years, I've just left Google Deepmind to start an AI Safety nonprofit, around Scalable and Human Oversight (i.e. building stronger "judges")! 🇨🇦 And, I'll be at FAccT in Montreal this week! (1/4🧵)
Turns out you can invert the Bellman equation to recover an agent's world model from its value function. Excited by the potential applications of this work, lead by @_aletcher. My fave bit - RL agents implicitly model latent variables they were never trained to optimize for..🧵
Jon Richens retweeted
I'll very soon be hiring a postdoc at UCL for a 2y project combining interpretability methods with behavioural evals to study whether goal and belief representations can be reliably extracted and manipulated in LM agents.
More details soon. Please get in touch if interested.
Theory + alignment ❤️. Really excited to see what comes out of this new org.
Jon Richens retweeted
🚨Transformers don't learn Newton's laws? They learn Kepler's laws!
Like us, transformers don't predict a flying ball via a differential equation, but by fitting a curve.
Moreover, reducing context length steers a transformer from Keplerian to Newtonian. Compression in play.
Jon Richens retweeted
Thrilled to share our new #NeurIPS2025 paper done at @GoogleDeepMind, Plasticity as the Mirror of Empowerment
We prove every agent faces a trade-off between its capacity to adapt (plasticity) and its capacity to steer (empowerment)
Paper: david-abel.github.io/plastic…
🧵🧵🧵👇
Jon Richens retweeted
I will be a SPAR mentor this Fall🤖
Check out the programme and apply by 20 August to work with me on formalising and/or measuring and/or intervening on goal-directed behaviour in AI agents
More info on potential projects here 🧵
Jon Richens retweeted
2 years ago, @ilyasut made a bold prediction that large neural networks are learning world models through text.
Recently, a new paper by @GoogleDeepMind provided a compelling insight to this idea. They found that if an AI agent can tackle complex, long-horizon tasks, it must have learned an internal world model—and we can even extract it just by observing the agent's behavior.
I wrote a blog post unpacking this groundbreaking paper and what it means for the future of AGI 👇
richardcsuwandi.github.io/bl…
Jon Richens retweeted
Can we trust a black-box system, when all we know is its past behaviour? 🤖🤔
In a new #ICML2025 paper we derive fundamental bounds on the predictability of black-box agents. This is a critical question for #AgentSafety. 🧵
Causality. In previous work we showed a causal world model is needed for robustness. It turns out you don’t need as much causal knowledge of the environment for task generalization. There is a causal hierarchy, but for agency and agent capabilities, rather than inference!
… and many more! Check out our paper arxiv.org/pdf/2506.01622, or come chat to me at #ICML2025. Joint work @GoogleDeepMind with @dabelcs, @alexis_bellot_, @tom4everitt