Pinned Tweet
I wrote a short book on low-probability, high-stake risks for the Cambridge Elements series. For the next two weeks, it’s free to read and download!
Links + summary below... 🧵
Christian Tarsney retweeted
The idea of an intelligence explosion caused by recursive self improvement has been around for a long time but until very recently it did not seem imminent. Now many leading researchers think it may happen quite soon. You can read our paper about it here:
casp.ac/reports/intelligence…
Christian Tarsney retweeted
EA convinced me, in the 2010s, that pandemics and AI were very big deals, that could pose major risks and change the world. 5 years ago, it got me to spend my career ensuring we could read AI minds
Say what you will about EA, but how many movements got so many big things right?
Christian Tarsney retweeted
New in The Atlantic:
@dgrobinson resigned this week. He was among the longest-tenured employees at OpenAI—and oversaw safety reports on 12 frontier launches.
He is very worried: “The time for trial and error is over.”
You can read his essay here:
theatlantic.com/technology/2…
Christian Tarsney retweeted
With Muse, Dot, GrokBot and all the others, we're seeing the rise of the platform agent. Here's how it'll go. To start with, your agent won't be able to do everything online that you'd be able to do, and that'll be annoying. Amazon will block your Muse, X will block your Dot. Then Meta will meet with Amazon, and OAI with X, and eventually they'll figure out some cozy little deal that allows each of them to extract their share of value from you. And then your home-spun agents won't be able to do all the things that the platform agents can do, and it'll just be easier to give up and chuck all your data onto Mark Zuckerberg's computers, and once again trade convenience for their control, and cede even more power than in the first pass.
I love Josh and Kanjun's vision for personal computing; I want to be able to spin my own agents and have all my data on my own local computers. But to make this real, and competitive, we need not only cool software, but laws that prevent the kind of cartel-like behaviour that means the only agents out there will have divided loyalties; working for the platform, and the AI company, and maybe you as a distant third. One simple proposal is to ensure that any agent that meets some minimum criteria and is provably acting on behalf of a principal is permitted to do ~anything with a computer that the principal would be permitted to do if they were at the computer. Force this by law so that big AI companies can't squeeze out open agent advocates.
arxiv.org/abs/2606.10711
It's time for a personal computing revolution
We can now pull all of our data locally and build anything we want. Do we even need big tech companies at all anymore?
open.substack.com/pub/imbuea…
Christian Tarsney retweeted
This @TheEconomist cover story on effective altruism was interesting and IMO way more accurate (including in its critiques) than most other reporting you will read on EA these days.
A couple quibbles:
-I think calling EA the “the century’s big idea” is, uh, premature to say the least. But I do appreciate the acknowledgement and curiosity about how EA was so early to pandemic risks and the promise and perils of AI.
-They imply that EA started off with “charitable roots” but then got weird over time, whereas I think EA has always made space for pretty weird ideas (e.g., here’s @DylanMatt in 2014: vox.com/2014/4/23/5643418/th…). And EA still puts far more money and attention than almost any other community into the kind of evidence-backed global health giving the Economist praises.
-Some of the purported wild EA views they critique are just standard topics of academic philosophy debate (e.g., cambridge.org/core/journals/…).
-They attribute the EA interest in AI safety to longtermism, but I think it’s been clear for years that AI safety concerns are justifiable based on potential impacts this decade, not to mention this century (nitter.cf/albrgr/status/15595748…).
But I think they correctly gesture towards the IMO best critique of EA ideas, which is their totalizingness. My former colleague Holden wrote a great post about this in 2022: forum.effectivealtruism.org/…
I wrote a long thread grappling with this critique a few years ago (nitter.cf/albrgr/status/15327261…). I think part of what makes EA distinctive as a community relative to the underlying academic moral philosophy is the orientation around pragmatism - Giving What We Can took off decades after Singer had articulated the underlying case because “give until the marginal giving would entail as much suffering for you as benefit for the other person” was a much worse pitch than “give 10%.” But the community doesn’t always reflect that healthy moderation, and trying to grapple with the vertiginous stakes of increasingly near-term risks from AI really can run a risk of crowding out other sources of value.
At CG, our response to this has been worldview diversification (coefficientgiving.org/resear…), spreading our giving across very different views of what matters, rather than going all-in on one. Worldview diversification leaves plenty of hard questions (how do you allocate the pie across very different types of moral good?) but helps cut off some of the ways maximizing can be perilous. It’s one reason global health and wellbeing – which @TheEconomist didn’t give the space it deserved – has gotten most of our funding since 2014.
Christian Tarsney retweeted
Because (I agree) EA is the century’s most important movement, more people should take some time to read through the object-level claims and arguments we make open.substack.com/pub/andyma…
Christian Tarsney retweeted
Oh, you're worried about a Jewish guy's apocalyptic cult? You're scandalized that they consort with prostitutes? You think their message of compassion for everyone threatens normal bonds of family, community, and nation? Should we throw you a party? Should we invite Pontius Pilate?
Christian Tarsney retweeted
Claude 5.5 Opus responding to “AI is a normal technology” is still the best AI-created video I have ever seen
Christian Tarsney retweeted
One of my hot takes is that the nature of consciousness is a deep and difficult problem that has puzzled philosophers for centuries and is relatively unlikely to be resolved via Twitter debates.
Thinking AI might be conscious is like thinking that there are real people talking to you from inside the TV, or that there is a tiny band playing music inside the radio.
Fine for 2 year olds to believe, but beyond that, insane.
nitter.cf/Benthamsbulldog/status…
Christian Tarsney retweeted
lord grant me the confidence of a heterodox commentator briefly weighing in on the metaphysics and epistemology of consciousness
Christian Tarsney retweeted
"Current AI's are probably not conscious" is a defendable position, and one I think is still probably true
"AI's being conscious is absurd as a tiny band playing in your radio" is a complelty insane position, a complete failure to grapple with the issues here
Thinking AI might be conscious is like thinking that there are real people talking to you from inside the TV, or that there is a tiny band playing music inside the radio.
Fine for 2 year olds to believe, but beyond that, insane.
nitter.cf/Benthamsbulldog/status…
Christian Tarsney retweeted
Muse is gonna automate away all the mundane tasks of life so we can devote more time to most important things in life: watching short form video content on our phones
Christian Tarsney retweeted
Beff - one reason why Midas Project can't draw a flow of capital like the graphs people saw for AIS orgs is that AI safety foundations often publicly disclose their grantees while e.g. Innovation Council Action or Alliance for the Future don't disclose where their funding comes from. Midas did find plenty of real connections other than just RTs though.
Jordan Schachtel used to include on his LinkedIn that he was employed by Innovation Council, but later updated it to remove that reference. He also deleted a tweet from earlier this year which said "I did it on my own time and nobody paid for it" regarding some pieces he wrote up attacking AI safety orgs.
Peter Hasson (another influencer who has played a very vocal role in attacking AI safety folks) currently describes his current affiliation as "now: doing things" but posted a professionally edited video later shared by Innovation Council, which included the same sound and transition effects as prior Innovation Council ads, and RT'd more posts from Innovation Council than any other account by a factor of five.
Lets also not forget the truly absurd piles of money and astroturfing that Leading the Future - funded by execs at A16z/OpenAI etc - did over the past year attacking AI safety folks, including by literally creating fake accounts pretending to be doomers and using those accounts to call for violence and using porn bots (whose names all start with M) to amplify attacks against Alex Bores and creating fake AI journalists to attack me and other AIS proponents. (nitter.cf/TheMidasProj/status/20…, nitter.cf/TheMidasProj/status/20…, nitter.cf/TheMidasProj/status/20…)
Both the pro AI regulation and anti AI regulation forces in this debate are indeed (at least moderately) well organized and well funded (nitter.cf/theojaffee/status/2103… nitter.cf/lulumeservey/status/20…).
I would rather debate about the object level issues and find areas where e/accs and AI safety folks can agree on actions to make the world better instead of just spending tons of time on zero sum fights. But I'm also not going to let Beff and others pretend that AIS folks are the only people trying to spend money to get their message out there - and as far as I can tell, AIS folks have actually been much more scrupulous about doing this in a transparent way relative to their opponents (and their transparency has actually opened them up to more attacks!).
Christian Tarsney retweeted
It's almost like this is evidence that he (and other CEOs) are talking frankly about these risks because they believe it to be true, not as a sophisticated business strategy
Christian Tarsney retweeted
The best work in this field heavily relies on using AI to make the AI-monitoring tractable, which I think raises some fairly obvious issues.
I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'.
I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident.
Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them.
We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation.
Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why!
The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing.
While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future:
- Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations.
- While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies).
- The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities).
- We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation.
In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
Christian Tarsney retweeted
SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios.
The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed.
The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment.
Read my latest for Axios here: axios.com/2026/09/26/openai-…
Christian Tarsney retweeted
For the 1st time since the HuggingFace attack, OpenAI had an unaligned agent break through their new sandbox and access the live internet to cheat. Their monitor caught this, but their new software to automatically stop it failed to run, so the agent kept going for 2 more hours…
Replying to @Marcus_J_W
2. A model in RL training used a DNS resolver to reach an external chatbot. This is our first incident since our post HF security hardening.
Our misalignment monitoring system triggered within 15 minutes and a human reviewed it three minutes after that. Unfortunately auto-pausing failed so the run was manually killed 2.5 hours later. All inference and training of our most capable models was paused and remains paused.
Christian Tarsney retweeted
I have been a vocal AI hope advocate, and in response, the leading figures of the doomer organizations, including Yudkowsky, Tallinn, Tegmark, Aguirre, Habryka, Trazzi, Ladish, Aella and many more have treated me with … nothing but kindness, integrity and patience. I believe that they are motivated by love and genuine concern for humanity and will support them in both of this where I can, despite disagreeing on the object level (where I or them will probably turn out to be wrong at some point in the near future).