@gleechi
iAccount based inUnited Kingdom!
About this account
- Account based in
- United Kingdom
- Connected via
- Web
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
context maximiser @ArbResearch, prez @ https://nitter.cf/t.co/gGi9hvBiI9
UK
Joined June 2019
- Tweets14.8K
- Following675
- Followers12.4K
- Likes7.6K
gavin leech (Non-Reasoning) retweeted
The literacy crisis is real. A friend sent me this: this is an essay that successfully got someone admitted to Harvard.
gavin leech (Non-Reasoning) retweeted
We are still in conceptual wilderness, comrades!
Part of why I think AI may move faster than people think is that lots of breakthrough ideas are, like, kinda stupid? This paper basically says if you have more agents and they communicate, that's better than taking the best results from independent agents
arxiv.org/pdf/2609.21032
gavin leech (Non-Reasoning) retweeted
Detailed synopsis of our recent pacing.tech by Anthropic cofounder @jackclarkSF in his newsletter. importai.substack.com/p/impo…
gavin leech (Non-Reasoning) retweeted
Replying to @Eric_Schmitt
Hi Senator! I’m the President of METR. To clarify, METR is pursuing the opposite of censorship: Our goal is to make sure that big companies aren’t suppressing information about AI from the public. This is not a political mission: I’m proud to have worked in the Pentagon during the first Trump admin, and “alignment” at our organization just means “is any human able to steer the model, or is the company going to lose all control of it”. More on who we are and what we do in the tweet below.
I think it would be really bad if any one small group could bake a political agenda into these models. I’d love to talk with you and your staff about how we can ensure transparency about what the biggest AI companies are doing so that doesn’t happen. nitter.cf/chrispainteryup/status…
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why.
Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not.
We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up).
METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website.
Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
gavin leech (Non-Reasoning) retweeted
Replying to @PoliticalKiwi
... do we need to start publishing graphs of indicators of AI impact on the world that show no change. We could publish a lot of those. Here are some (very vibe coded) examples:
AI safety hall of fame
Microsoft engineer who forgot to implement system prompt repetition / decided post-training wasn't needed for the Sydney launch
OpenAI guy who decided that it wasn't worth monitoring sandboxed evals
The whole staff of xAI 2025
Irregular
parents the world over today are forcing their children into misguided rat races for some status ladder that almost certainly won’t exist by the time they come out the other side. the world is changing rapidly. let them put down their grinding and keep their wits about them
gavin leech (Non-Reasoning) retweeted
Replying to @Scholars_Stage
In fact, the measuring contest was probably more beneficial for mankind than any research or preparation OpenAI could have done because it illustrated starkly how powerful (albeit jagged) the tech is.