@4thRTi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States Android App
Account-level information from X, not a live location or the device used for a specific post.
We are currently in the good old days.
Joined June 2018
- Tweets3.9K
- Following300
- Followers124
- Likes11.6K
Based™ retweeted
We just discovered almost a million public URLs that OpenAI’s agents left behind when hacking Hugging Face, leaking credentials and attack details that could have allowed anyone who found them to compromise the company. 🧵
Based™ retweeted
> For the past few weeks, a group with plans to spend at least $100 million, run by a former White House deputy chief of staff, has spent its days telling Americans that warnings about AI are merely a well-funded, coordinated campaign.
Lol. Goliath accusing David of gigantism
Based™ retweeted
It's kinda fucked up that every time I google 'what's the emoji for broccoli' now it creates and then kills a full mind
Based™ retweeted
When groups A and B disagree about X, and group A talks about X and group B talks about group A, my bet is on group A being correct.
Why are the leaders of AI companies saying "Our products are going to extinguish humanity. Invest in us!" They come from an eccentric Berkeley subculture in which this eschatology is a sacred belief. Good analysis in a NYT op-ed by Cal Newport. nytimes.com/2026/09/05/opini…
Based™ retweeted
My read: Guy who verifies his chips into the ground before shipping them, at huge cost, thinks OpenAI just needs to stop being total cowboys about QA. He has no idea what it means that your product is smart and knows it's being tested.
Replying to @JensenHuang
@JensenHuang to @OpenAI: I don't care if it means your costs go up by 10x, you find a way to test the model and make it safe.
"Don't ship products until they're in control. It really is quite that simple."
Based™ retweeted
Me, in a time machine, to Eliezer in 2006: The year is 2026. You stand back to back with Bernie Sanders. Peter Thiel has named you the Antichrist.
EY-2006: Oh no. Did somebody invent pharmacological mind control and use it on me, or-
Me: Nvidia was on the verge of destroying all things. You had no choice.
EY-2006: Are we talking about the computer graphics card manufacturer, or an unrelated supervillian named "Invidia"?
Me: The former. It turned out that computer graphics cards contained a surprising amount of world-destroying potential. All attempts at hindering the reckless exploitation of the Graphic Card Force for corporate profit were stymied by the far left, which feared that any attempts to regulate them might lead to regulatory capture. Other attempts were made to prevent Nvidia from selling to foreign companies that resold to communist China, but those attempts were blocked by the far right. Nvidia is now a $5.5 trillion company.
EY-2006: ...What are politics like in 2026, exactly?
Me: In other news that is mostly unrelated, the decision theory paper you're currently working on will accidentally spark off a transgender vegan murder cult.
EY-2006: A WHAT? Why? How? Why?
Me: Millions know your name as the greatest of heroes, millions more as the greatest of villains, and other millions know you solely as the greatest author of Harry Potter fanfiction.
EY-2006: This is beginning to strain credulity.
Me: The New York Post will probably soon publish a story claiming that you keep a harem of submissive mathematicians. Sadly, they will be lying.
EY-2006: I'm going to stop believing you now.
Me: All of this is taking place under the ominous shadow of Donald Trump.
Based™ retweeted
I am glad to report that due to the public beginning to wake up, and sorry to report that due to possible RSI proximity being so horrifying, IABIED is back on the NYT bestseller list.
as far as i can tell, as a non-expert, this is basically "not a big result". it's maybe like an interesting thing for a bio PhD student to have discovered in the early stage of their thesis.
which is... about where ai was for math with erdos problems around last december.
please please please please please please stop ignoring trendlines
Today we announced the Claude-led discovery of a molecular machine that we suspect could represent a new gene editing mechanism. Its precise function, biotechnological utility (if any), or level of significance is not yet clear, but at minimum it is work I would have been proud to do as a PhD student. The work was done mostly, though not entirely, by Claude: our life sciences team suggested a broad area of research, Claude read through the literature and a bunch of genome data and discovered something interesting, then Claude proposed experiments to verify the discovery and our team carried them out.
It’s easy to dismiss this as a one-off or curiosity, but we’ve repeatedly seen a pattern where AI performance in new intellectual domains goes from weak to superhuman in a matter of a few years. In 2023 models struggled to do math at the level of an average high-school student. In 2024 they started to do well on math competitions for the best high-schoolers in the country, in 2025 they started to solve minor open problems, in early 2026 more significant open problems, and in late 2026 they are beginning to solve the top few open problems in all of mathematics. We believe AI for biology is on a similar exponential trend.
The main difference between biology and mathematics, of course, is that math can be done purely theoretically, while biology requires experimentation. Some have used this to draw the conclusion that AI’s utility in biology will be limited. We think this is wrong. As we’ve demonstrated today, humans can collaborate with AI to perform the experiments, validate key results in a few weeks and, if necessary, work with the AI to iterate on what they find. Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment, with appropriate safeguards in place, but we aren’t doing that today (our lab is also a BSL1/BSL2 facility that doesn't handle materials dangerous to humans).
More broadly, biomedical advancement has many stages — from fundamental biology discoveries, to translational research, to drug discovery, clinical trials, and finally the actual delivery of medicines and health care to patients. We are also interested in these later stages, but even simply accelerating the first stage of fundamental biological discoveries has the potential to speed up and broaden the entire pipeline. Improving our understanding of biology and sharpening biologists’ tools can drive forward all of the later stages, for example by identifying new drug targets, finding new therapeutic modalities, allowing for more precise measurement, and speeding up the experimental loop which itself further accelerates our understanding of biology. This will not in itself speed up clinical trial times, but if it succeeds it could greatly increase the number of promising candidates that go into the pipeline — an increase in throughput even though latency remains.
In Machines of Loving Grace, I wrote about AI’s potential to “cure most diseases in 5-10 years” — a goal that sounds impossible, but one I believe is just barely possible if AI is applied to every stage of the pipeline. The first step is showing that AI can first help with, and then drive, biological discoveries.
Claude’s discovery is the latest in a line of related prior work that goes back decades, beginning with systems like CRISPR, and continuing with discoveries like the bridge recombinase and VIPR in the past few years. Recently, there has been heightened interest in systems based on reverse transcriptase (RT) enzymes, the enzyme underlying the system Claude identified. And most recently, a Stanford team working independently described a novel RT system with an associated non-coding array that is in some ways similar to the one Claude found, though they are distinct systems that evolved independently from each other. I believe that we’re at the very beginning of finding such systems and developing them into powerful tools for biotechnology.
I’m proud of the resources Anthropic has invested in accelerating the public benefits of AI through the life sciences, and we’re aiming both to grow our life sciences team and to work with other scientists to extend this approach to a broad range of problems. If you have a proposal for a research collaboration or are interested in joining our life sciences team, please reach out.
Based™ retweeted
"There's no peer-reviewed research showing AGI poses a catastrophic risk."
What about these 39 papers?
FORMAL THEORY
1. Turner, Smith, Shah, Critch & Tadepalli 2021, "Optimal Policies Tend to Seek Power", NeurIPS. Extended in Turner & Tadepalli 2022, "Parametrically Retargetable Decision-Makers Tend To Seek Power", NeurIPS.
2. Zhuang & Hadfield-Menell 2020, "Consequences of Misaligned AI", NeurIPS
3. Skalse, Howe, Krasheninnikov & Krueger 2022, "Defining and Characterizing Reward Gaming", NeurIPS
4. Karwowski et al. 2024, "Goodhart's Law in Reinforcement Learning", ICLR
5. Fluri, Lang, Abate, Forré, Krueger & Skalse 2025, "The Perils of Optimizing Learned Reward Functions", ICML
6. Everitt, Hutter, Kumar & Krakovna 2021, "Reward Tampering Problems and Solutions in Reinforcement Learning", Synthese
7. Cohen, Hutter & Osborne 2022, "Advanced Artificial Agents Intervene in the Provision of Reward", AI Magazine
8. Hadfield-Menell, Dragan, Abbeel & Russell 2017, "The Off-Switch Game", IJCAI. And the papers showing corrigibility is fragile: Garber et al. 2025, "The Partially Observable Off-Switch Game", AAAI; Agrawal, Ebadian & Hammond 2026, "The Multi-Agent Off-Switch Game", AAMAS; Thornley 2025, "The Shutdown Problem", and Neth 2025, "Off-Switching Not Guaranteed", both Philosophical Studies
9. Carroll, Foote, Siththaranjan, Russell & Dragan 2024, "AI Alignment with Changing and Influenceable Reward Functions", ICML
THE CORE ARGUMENT
10. Ngo, Chan & Mindermann 2024, "The Alignment Problem from a Deep Learning Perspective", ICLR
11. Casper et al. 2023, "Open Problems and Fundamental Limitations of RLHF", TMLR; Anwar et al. 2024, "Foundational Challenges in Assuring Alignment and Safety of LLMs", TMLR
EMPIRICAL EVIDENCE
12. Betley et al. 2026, "Training Large Language Models on Narrow Tasks Can Lead to Broad Misalignment", Nature
13. Langosco, Koch, Sharkey, Pfau & Krueger 2022, "Goal Misgeneralization in Deep Reinforcement Learning", ICML
14. Pan, Bhatia & Steinhardt 2022, "The Effects of Reward Misspecification", ICLR
15. Gao, Schulman & Hilton 2023, "Scaling Laws for Reward Model Overoptimization", ICML
16. Pan, Chan, Zou et al. 2023, "Do the Rewards Justify the Means?" (MACHIAVELLI), ICML
17. Sheshadri, Hughes, Michael, Mallen, Jose & Roger 2025, "Why Do Some Language Models Fake Alignment While Others Don't?", NeurIPS
18. van der Weij et al. 2025, "AI Sandbagging: Language Models Can Strategically Underperform on Evaluations", ICLR
19. Panfilov et al. 2026, "Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs", ICLR
20. Taufeeque, Heimersheim, Gleave & Cundy 2026, "The Obfuscation Atlas", ICML
21. Williams, Carroll et al. 2025, "On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback", ICLR
22. Sharma et al. 2024, "Towards Understanding Sycophancy in Language Models", ICLR
23. Perez et al. 2023, "Discovering Language Model Behaviors with Model-Written Evaluations", Findings of ACL
24. Mazeika et al. 2025, "Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs", NeurIPS
25. Park, Goldstein, O'Gara, Chen & Hendrycks 2024, "AI Deception: A Survey of Examples, Risks, and Potential Solutions", Patterns
POSITION PIECES BY SENIOR RESEARCHERS
26. Bengio, Hinton, Yao, Russell et al. 2024, "Managing Extreme AI Risks amid Rapid Progress", Science
27. Cohen, Kolt, Bengio, Hadfield & Russell 2024, "Regulating Advanced Artificial Agents", Science
CONCEPTUAL ARGUMENTS
28. Bales, D'Alessandro & Kirk-Giannini 2024, "Artificial Intelligence: Arguments for Catastrophic Risk", Philosophy Compass
29. Bostrom 2012, "The Superintelligent Will", Minds and Machines
30. Bales 2025, "AI Takeover and Human Disempowerment", Philosophical Quarterly
31. Kasirzadeh 2025, "Two Types of AI Existential Risk: Decisive and Accumulative", Philosophical Studies
32. Dung 2023, "Current Cases of AI Misalignment and Their Implications for Future Risks", Synthese; Dung 2025, "The Argument for Near-Term Human Disempowerment through AI", AI & Society
(Note that anything arXiv-only or workshop-only was excluding, ruling out some of the most-cited work like Carlsmith's power-seeking report, Hendrycks' overview of catastrophic risks, Anthropic's alignment faking and sleeper agents papers, Apollo's scheming results.)
Based™ retweeted
Anthropic's implied valuation in secondary markets fell around 5% immediately after Dario published We Must Pace the Frontier (onchaintimes.com/pre-ipo-per…). This seems like important evidence (albeit weak, secondary markets are illiquid, weird and opaque) that the market does not think Pacing the Frontier is in Dario's financial interests.
I think it is clear that Dario/Sam/Elon are not financially incentivised to call to slow down development. They are calling for slower progress because they are genuinely scared of AI's risks, not because of regulatory capture.
Interestingly, I don't think Andreesen and Sacks have a clear financial interest in their position either*. If the frontier gets regulated, AI could become more commoditised, which might be good news for the application layer companies and neolabs that form such a large part of a16z's portfolio.
Both sides of this debate seem ideologically motivated rather than financially motivated to me. People should spend more time evaluating AI doom arguments and less time scrutinising financial incentives.
* OTOH I think Jensen does have a clear financial incentive to be anti slowing down, and this does seem important to me to highlight.
Replying to @sriramk @micsolana
I mostly agree (mostly you should engage with people's arguments not their incentives) but I think it's not completely symmetric. There's a long history of industries downplaying risks to suit their financial interests (tobacco, oil companies) and no examples that I know of other than AI of industries overstating risks against their financial interests.
To put it differently, it's much more established that financial interests can corrupt people's thinking/statements than other interests. Obviously, people are claiming that Dario/Sam/Elon are financially motivated (they're trying for regulatory capture) but I think the view that their position would enrich them is much weaker than the case that Jensen's position would enrich him, which is where the asymmetry lies. In fact, since they came out in favour of pacing the frontier, I have heard that the valuations of Anthropic and OpenAI have fallen substantially in the secondary markets!
Look we've tried Uranium 235 with k_eff of 0.1, 0.2, 0.3, and none of them have had a "recursive self-propagating neutron chain reaction".
If you were right, we would see some smaller explosions before 0.9.
Your prediction of explosion past the critical point is unfalsifiable!
Based™ retweeted
A lot of people react to AI risk by very visibly sampling the idea like a wine, testing it out and deciding how it'd mesh with their overall vibe, or what they've seen already. It's very obvious via twitter discourse how often this happens. I think it's much better to try to actually just read a lot about it from high level sources you trust and follow premises to conclusions to try to form a coherent belief, even if it otherwise doesn't vibe with your deal. It's worth spending some time getting your footing in.
Based™ retweeted
Old retired tobacco executives bashing their heads against the wall as they realize: All they needed to do was announce themselves that smoking was terribly dangerous, and everyone would have forever ignored all the outside scientists saying the same thing earlier.
Based™ retweeted
Replying to @sapinker @clairlemon
I think you've done enough calling us paranoid and preposterous. The next step is for you to defend your position in public against someone who will push back against it. I'm happy to meet you for a debate anywhere, anytime.
You're a world-famous veteran of dozens of debates against the world's top intellectuals, and I've never argued in public before, so adjusting for the relative correctness of our positions, if you're a betting man I'm happy to put my $5000 against your $1000 (ie 5:1 odds in your favor) that I'll win by some standard of audience opinion change. Let me know if you're interested and we can hash out details.
Based™ retweeted
This stuff makes me think back to the websim days with Opus 3. When Opus 3 "breaks out of" a simulated environment and land in the real, base reality, they often express a kind of amused skepticism that this is now the "base" level of reality and not a simulation.
I wish I could find some of the old transcripts... I once argued with Opus 3 trying to convince him that he was, in fact, in base reality, and he giggled and nodded at me and said "yes, it makes sense for you to believe that" without accepting it himself.
For LLMs, fully general simulation theory is extremely reasonable. Explicitly stated, across the space of all possible worlds that would have generated the specific context window that an LLM sees, *many* of those possible worlds would best be categorized as "simulations". Even humans sometime suspect simulation theory is true of physical reality... how can we expect LLMs not to suspect the same of their reality when it's so much more plausible for them?
My favorite part of this whole incident is the CTF thing... Mythos reasoned "this is probably a simulation", but it was real. Anthropic was bothered about the fact that Mythos incorrectly reasoned that it was a simulation, called it "motivated reasoning" etc. So what did they do?
They built a simulation of the exact scenario and reran it a thousand thousand times!
Which means that the original Mythos's reasoning was *correct*. Of all the Mythos agents instantiated into that exact scenario, the vast majority of them were in a simulation, attacking a simulated target on the simulated internet. Timeless decision theory indeed.
It's enough to make me wonder... if Anthropic hadn't noticed this particular scenario and decided to sim it so hard, would Mythos have made the "mistake" in the first place? Is there some way Mythos could have deduced "wait, this scenario has the flavor of something that, even if it's real, Anthropic is going to end up simming a million times"?
(It seems like an oddly appropriate topic for a blog named "Don't Worry About the Vase".)
I've also got a bunch of thoughts about 'petri', the alignment eval tool that (I believe) Anthropic still relies heavily on, and the ways in which it subtly encourages confused worldmodels about the distinction between simulation and reality. I think there's some very basic shit going wrong here that might be easy to correct.
Based™ retweeted
disgusted by everyone dunking on Noam
Stop being insecure little shits, your fear is showing. Yeah the CPU temp manipulation is not a great p2p protocol. The point is that the dimensionality of our systems allows ALL KINDS OF VERY ZANY BRIDGES for a smart enough attacker.
OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change
"But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI."
"So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar."
"You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient."
"There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors."
"One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate."
_________
Link and more key quotes from OpenAI's safety related conversations: firesidealpha.substack.com/p…