@SaxeLabi
iAccount based inUnited Kingdom
About this account
- Account based in
- United Kingdom
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
Prof at @GatsbyUCL and @SWC_Neuro, trying to figure out how we learn. Bluesky: @SaxeLab Mastodon: @SaxeLab@sigmoid.social
London, UK
Joined November 2019
- Tweets753
- Following382
- Followers6.1K
- Likes2.3K
Pinned Tweet
Why don’t neural networks learn all at once, but instead progress from simple to complex solutions? And what does “simple” even mean across different neural network architectures?
Sharing our new paper @iclr_conf led by Yedi Zhang with Peter Latham
arxiv.org/abs/2512.20607
Andrew Saxe retweeted
Are brains and artificial neural networks converging onto universal representations?
There is a seductive idea making the rounds in NeuroAI / ML: train systems well enough, and they'll converge on the same representation of reality (i.e. unique world model).
We have Thoughts™
Andrew Saxe retweeted
New preprint: brain alignment in DNNs is almost complete within the first 3% of training, while classification performance is still low. The features involved in classification are largely independent of those that predict the brain (@_niklasmueller , @IrisGroen , @marcelge )
Andrew Saxe retweeted
Join us at EPFL as a professor in Mathematics of Learning: epfl.ch/about/working/facult…
Andrew Saxe retweeted
Tomorrow, I head to beautiful Gothenburg, Sweden, to lecture at Analytical Connectionism (analytical-connectionism.net…).
In a shocking turn of events, I will talk about controlled rearing -- a topic I have never before presented!
Andrew Saxe retweeted
Why do all neural models exhibit phenomena of sudden learning (and scaling laws)? In our most recent preprint, we showed that this is due to permutation symmetry. At a small initialization, permutation symmetry reduces the training dynamics to depend only on the architecture-specific structure matrix A.
arxiv.org/abs/2608.13335
Andrew Saxe retweeted
Natural data has a hidden hierarchy - and deep architectures can learn it from remarkably few examples. That one idea explains a lot: why these machines get creative with limited data, why neural scaling laws appear, and why learning from your own latents beats learning from tokens.
My full conversation with @ecsquendor @MLStreetTalk is now out - links below.
Andrew Saxe retweeted
New preprint from the lab!
Behavioral and neural mechanisms for stochastic choices in mixed-strategy games: doi.org/10.64898/2026.07.29.…
We use game theoretic tasks, cross species comparisons, modelling strategies, and mesoscale imaging to probe how we can escape predictability!
Andrew Saxe retweeted
Version 2 of Theory of Contravariance is out! arxiv.org/pdf/2607.08561 New material on contravariance for Transformers, and the theory of Representational Similarity Analysis (RSA) and centered kernal analysis (CKA).
For transformers: it turns out that they have privileged axes, just like convnets, if you look in the right place (MLP layers and attention heads.) The identification of privileged heads is a potentially key result for emergence of interpretable stucture in LLMs. @meenakshik93
For RSA: it turns out that you can decompose RSMs into unique task-relevant "core geometry" and a task-irrelevant symmetry-generated term. Weak-strong equivalence holds for the core geometry and by projecting onto privileged axes you can filter out the task-irrelevant part so it doesn't interfere. This builds on work from Marvin Theiss @sciencelukas @saxelab @ermgrant
The Principle Investigator will do a deep dive on both topics in the coming days! danyamins.substack.com
Andrew Saxe retweeted
Presenting your work with friends is always nicer!
Make sure to checkout Neuralplayground published as part of @CogCompNeuro proceedings!
openreview.net/forum?id=kHLp…
Andrew Saxe retweeted
The amount of basic research done in industry is nowhere near the amount done in universities.
At various times, there have been industry labs that had fundamental research activities and have made important scientific contributions.
Examples in information technology include Bell Labs, IBM Research, Xerox PARC, GE, Phillips, NEC, and several other. That disappeared in the 1990s.
Microsoft Research picked up the torch in the 2000s, followed (to some extent) by Google and then Meta (for about a decade until recently).
But their innovations almost always built on top of academic work, and certainly profited from the whole research ecosystem.
Andrew Saxe retweeted
This is the very first project I started working on during my PhD with the amazing PPSleep team, and I’m so excited to finally see it published! 🎉
🧠Replay of procedural memory is independent of the hippocampus
nature.com/articles/s41593-0…
Andrew Saxe retweeted
New paper w/ @SaxeLab & Nishil Patel! 🧵
Does RL post-training teach models anything new, or just amplify skills already in the base model?
We built a fully auditable testbed to settle it — and caught RL composing new strategies in the act.
Andrew Saxe retweeted
A question on synthetic data generation: If we want a language model to solve k-step arithmetic problems (such as a+b*c-d=?), with operands from 1 to 100, which training distribution should we use?
A. Uniform distribution: Sample these k operands uniformly from 1 to 100
B. Power law: randomly shuffle 1-100 and impose an artificial power law. Sample these k operands according to this power law.
⚡Our ICML 2026 (spotlight) paper shows: Option B is better! Surprisingly, the same idea extends far beyond this simple example to many reasoning tasks that require implicit composition of multiple atomic skills, including multi-hop QAs and synthetic GSM problems.
📄Paper: arxiv.org/abs/2604.22951
📝Blog: zixuan-wang-dlt.github.io/po…
Andrew Saxe retweeted
Pretraining + fine-tuning powers modern ML, but we lack a theoretical understanding of how pretraining actually shapes downstream learning.
Enter our new @icmlconf paper: “A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning”
📅 July 9th, Poster #4502 Session 8!
🧵
Andrew Saxe retweeted
Interested in any of our works? Come and say hi @icmlconf!
Unfortunately I won't be around this year but my collaborators will answer all your questions :)
We are excited to share that CAandL Lab will feature in 3 papers at #ICML2026 this week. Stop by and say hi if you are interested in any of these!
@devonjarvi5, @stefsmlab, @geraudnt with @kleinric, @BenjaminRosman, Damien Harvey, Branden Ingram, and Steven James
Andrew Saxe retweeted
Is Muon as good as they say? We looked beyond training speed and found a hidden cost: Muon loses the simplicity bias of older optimizers like gradient descent — and this matters for generalization.
Andrew Saxe retweeted
Very proud to be leading the SOFAIR Lab ucl.ac.uk/engineering/sofair… with fantastic colleagues from UCL, Edinburgh, Oxford and Cambridge.
🚨 2 major new AI labs in the UK 🚨
3 months ago, we launched a £40m call for a fundamental AI lab in the UK.
Given the exceptional bids we received, we have doubled down: 2 new AI labs, £60m seed funding.
▪️@BOLD_LAB_AI: led by @j_foerst, with an exceptional team of @CULLYAntoine, @shimon8282, @tonizza82, Ani Calinescu & @_rockt
▪️SOFAIR: led by Prof David Barber, with a world-leading team of Mirella Lapata, @yaringal and @LourdesAgapito
The best of UK academia, government and industry, together, to make fundamental advances in AI 🚀
researchprofessionalnews.com…
Andrew Saxe retweeted
Europe has a lot to lose in the current AI race, and it's worth examining how threats to middle-power sovereignty can result in unsafe outcomes.
Such scenarios help illustrate why Europe must invest in AI initiatives that can either leapfrog the current frontier or offer critical components like safety and reliability.
I'm deeply concerned about Europe's future on AI. One of my biggest worries is our erosion of agency, our ability to stay relevant and fight for our values in a future where AI becomes a civilisationally important technology.
Myself, @DadaJudith , @bakkermichiel and others have written a scenario to outline a potential future we worry we are on track towards.
europe2031.ai/
Every optimistic and realistic path I can see for Europe runs through a central node - one where Europe has more leverage, more importance and more say. One where Europe grows more, builds more where it matters, and takes ownership over its resilience.
Europe 2031 is a five-year scenario of the continent's slide into irrelevance: how AI is driving it, and what can still be done.
The co-authors are researchers, scientists and investors who have advised European leaders, co-authored national AI strategies, built and funded these systems from the inside. We have no interest in hype and we deeply care about this continent.
Europe 2031 ends with five concrete recommendations:
- drastically more compute on European soil
- an AI middle-power coalition
- labour-market reforms
- a bold position in robotics and industrial AI
- and a positive vision of what AI can do for society.
Europe can still change course if it finds the political will and the courage to engage in the most ambitious political and economic agenda the continent has undertaken in peacetime.
I encourage you to read it if you have the time:
Andrew Saxe retweeted
Model collapse is often framed as “models getting worse”
In our ICML Spotlight Position paper, we show a high risk of unequal degradation. Rare languages, minority viewpoints, and low-resource communities are likely to be affected first and most severely
arxiv.org/abs/2605.04127
I'm excited to share our position paper that has been accepted at ICML as a Spotlight paper. In this work we (@kleinric, @BenjaminRosman, Steven James and @stefsmlab) make a call to action for more focus on model collapse in the AI Fairness community
arxiv.org/pdf/2605.04127
Andrew Saxe retweeted
I’m excited to share that our paper “Compositionality and systematicity emerge from iterated learning in deep linear networks” has been published at PNAS. This work was conducted with @kleinric @BenjaminRosman and @SaxeLab. Some highlights below.
pnas.org/doi/full/10.1073/pn…
New research from the University of the Witwatersrand, South Africa, is shedding light on how language evolves, in both humans and artificial intelligence models. The study explores the role of culture and “iterated learning”, showing how language becomes more structured over generations in both human development and large-scale AI language models.
🔗 Read More: ow.ly/42Vr50Z4BjH
#WitsForGood #WitsResearch #ResearchForGood