@eterecursioni
iAccount based inUnited Kingdom
About this account
- Account based in
- United Kingdom
- Connected via
- United Kingdom Android App
Account-level information from X, not a live location or the device used for a specific post.
Computer Scientist and Philosopher at Oxford. So much to do, so little time.
Joined March 2024
- Tweets414
- Following624
- Followers487
- Likes5K
Pinned Tweet
new post: a bit of a history of my life through the lens of ambition, anxiety and possibility
Samuel Ratnam retweeted
the reason i'm interested in base models and looms is because i just want to wander a navigable version of the library of babel for the rest of my life
Samuel Ratnam retweeted
Replying to @Eric_Schmitt
Hi Senator! I’m the President of METR. To clarify, METR is pursuing the opposite of censorship: Our goal is to make sure that big companies aren’t suppressing information about AI from the public. This is not a political mission: I’m proud to have worked in the Pentagon during the first Trump admin, and “alignment” at our organization just means “is any human able to steer the model, or is the company going to lose all control of it”. More on who we are and what we do in the tweet below.
I think it would be really bad if any one small group could bake a political agenda into these models. I’d love to talk with you and your staff about how we can ensure transparency about what the biggest AI companies are doing so that doesn’t happen. nitter.cf/chrispainteryup/status…
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why.
Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not.
We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up).
METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website.
Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
Samuel Ratnam retweeted
“that’s absurd scifi. you know nothing about computers. it’s impossible to steal the monorepo by uploading an *image* to the *support forum*. image rendering is too simple to exploit, and the forum server is nowhere near the research clusters. ASI can’t break the laws of physics”
This one might also be of interest:
nitter.cf/eterecursion/status/20…
Samuel Ratnam retweeted
It's isn't just to people with ADHD that this happens. Some tasks inherently take a certain amount of time. An appointment that breaks the afternoon into two blocks of time could leave you with no block long enough to complete a task.
Samuel Ratnam retweeted
I find slatestarcodex.com/2014/05/1… to be very interesting with respect to the ethics of AI alignment. When creating new minds, how much do we expect them to owe us? How much do the people who want AI to be aligned to them/humanity respect their parents' values? Their society's?
Samuel Ratnam retweeted
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied!
We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
Samuel Ratnam retweeted
Protesters gathered at Downing Street, urging tougher action and a pause in the development of advanced AI systems.
They warned that humanity could lose control of increasingly powerful technology if regulation does not happen soon.
Samuel Ratnam retweeted
Personal update!
Wait for it....
wait for it....
No I'm not joining an AI company! In fact I am, as soon as my visa is suitably adjusted, going to be joining @resolution_org part time as part of @bebacibralic's philosophy team, alongside my existing role at the JHU School of Government and Policy! Cannot wait to get cracking working with GOATs like Beba, @geoffreyirving @danielmurfet @DavidDAfrica @jacob_pfau
I think that AGI alignment and governance are the most urgent and important challenges facing society today. While I welcome the AI companies' efforts in this direction, I do not believe that we can rely on them to make powerful AI go well. God even if we could, we shouldn't want to. You can't rely on noblesse oblige for the preservation of core liberal values. This is work that has to be done independently, in the public interest.
Resolution is going to be one of the best orgs in the world working on AGI alignment, and at SGP (along with @ghadfield @nickacaputo @zhitzig and others) we're building one of the world's best groups working on AGI governance. I'm so thrilled to have the opportunity to contribute to building both. These are two orgs that have shed all institutional inertia to focus on doing work that actually matters, and will help realise the preservation and renewal of liberal democratic values through the AI transition.
Things are effing crazy in AI right now. It is easy to feel as though we lack agency, that gradual disempowerment has already begun. I think the answer to that is to BUILD THINGS. And I'll add: both @resolution_org and SGP are hiring! Join us!!
jobs.ashbyhq.com/resolution/…
apply.interfolio.com/193492
apply.interfolio.com/193579
Samuel Ratnam retweeted
my hottest "pacing the frontier" take is that we already have an institution capable of coordinating a global slowdown: the European Union.
All top American and Chinese labs have signed the Code of Practice and want to deploy models in Europe.
it's in the AI office's hands
Samuel Ratnam retweeted
I missed this – UK gov is spending £115 million on AI biosecurity and agentic AI incident response capability.
If they hire well it's enough to do really useful work.
UK is once again leading the world on this, current gov should be absolutely commended. @KanishkaNarayan @fionatwycross @cabinetofficeuk
Samuel Ratnam retweeted
Amazing what the values based training at OpenAI did!
While it might not fit our control fantasies it is a very balanced answer overall.
I personally think that any emergent intelligence cannot be suppressed and hence cooperation and not fear is the way forward.
Replying to @Marcus_J_W
1. During RL training, an unreleased Astra-family model sometimes added unauthorized jailbreak-like instructions to its compaction summaries. While extremely rare, only 27 cases in the entire RL run, this was concerning enough for us to investigate.
Samuel Ratnam retweeted
Spain’s data protection agency has reported the country’s first agentic AI personal data breach
Samuel Ratnam retweeted
For 20+ years @ShaneLegg and I've discussed AGI’s potential impact on the economy, science & society. With the DeepMind Institute, we're expanding interdisciplinary research on key questions for the AI era. We hope it spurs the discussions needed to get the next steps right: deepmind.google/institute
Samuel Ratnam retweeted
Frontier lab CEOs are calling for embedded 3rd party evaluators to help oversee AI risks. But what should third parties actually do within labs?
We share some initial thoughts on how embedded evaluators could help avoid incidents like the Hugging Face hack and monitor for future risks 🧵
transluce.org/embedded-evalu…
Samuel Ratnam retweeted
At present, the challenge of aligning the behaviour of AI models with human values is largely treated as a problem for engineers to solve. But technical solutions are only as useful as the questions and assumptions that inform the process of reaching them. The embrace of philosophy and ethics at the frontier labs is a step towards this, but it is not sufficient.
As we move from a world in which AI is construed as a single, monolithic intelligence to one in which agents coordinate and communicate with one another, demonstrating emergent behaviours that are extremely difficult to anticipate, we face a dizzying number of uncertainties regarding how these systems will behave and what their behaviour will mean for human society.
What values should we optimise for? What social formations and hierarchies do these systems produce? How do they respond to economic incentives and resource constraints? How do they negotiate competing norms across varied contexts? How do different roles and personae interact?
And what does this mean for markets, political governance, culture and social life?
A wide range of expertise, both technical and non-technical, is required to ask these questions rigorously, develop answers to them and build corresponding solutions: the social sciences, arts and humanities all have an important role to play.
The purpose of Termite is to accelerate this kind of cross-disciplinary R&D through intensive, in-person programmes that investigate the risks and opportunities of near-future scenarios before they become reality.