@Raemon777

Secular Solstice guy

Berkeley
Joined August 2009
ASI risk skeptics: "How will ASI do the crazy science they'd need to invent novel superweapons? You keep saying 'it invents nanotech' but that sounds magic?" Rationalists: "Well look intelligence is very powerful and can be used to search the vast space of possible things and even though we can't predict exactly how Magnus Carlson is..." skeptics: "Sorry I feel asleep tl;dr. Shame how rationalists never have an answer to this question!" reality: "Anthropic is building the AIs a wetlab and... wait these guys are... straight up trying their hardest to build them a nanotech lab?" (I'm not actually sure whether a wetlab or this thing are more 'like nanotech' in the relevant sense." ... ...now, guys, this tweet here is representative of the bad part of twitter. I think the argument here is fair and the punchiness/comedic timing feels okay as an isolated drop in a bucket. But the aggregate of people posting Fun Dunking Takes on twitter seems bad for the world. I can feel Twitter making me a more tribal, worse person. I'm posting it anyway because... ...well, partly because it's funny, but I will endeavor to pay a sin tax of writing a thoughtful real argument here where I will try to meet people where they're at. Let's see if that helps. Many people are, correctly, very worried about stagnation. And about government overreach, or tyranny. Meanwhile all this "ASI might kill us" stuff feels fake. But, more importantly, the "and aligning ASI is so hard, the only options are on the order of 'global shutdown, buying decades'" feels _extra_ fake. Why can't we just do kinda normal regulation and iterate like we usually do? I think it's reasonably easy to understand "something smarter than humans we brewed in a lab might be dangerous." But, the arguments for "and alignment is probably VERY hard" are subtler. And I think it is super reasonable to have an immune reaction to arguments about the apocalypse that involve a number of subtle arguments you have to keep in your head at the same time. I won't try to bridge all that in this tweet. But: If you're the sort of guy arguing on twitter about whether ASI is safe, I think you have a responsibility to think about it. And that includes actually keeping track of when one argument is decisively answered. I'd *like* to hold you to "keep track of the evidence that accumulates for an argument being answered". This is harder. For, now, I'll settle for "the argument has been decisively answered." The answer to "how will the AIs get control of laboratories, and do all the real-world experiments that are necessary to do science?" is: "The humans will give it to them. The humans will actively try to give them full control over the process so the AIs can move faster and trim down the iteration time. Whatever the iteration time was with human researchers, it will become faster." This is what will happen, unless you stop people from doing that. Are there other bottlenecks and questions and considerations about exactly how this will play out, and how it weighs against other policy questions? Yes. But, keep track of the list of questions that are settled. Please. Other settled questions include "will AIs have motivation to deceive their overseers, and do excessive bad things the humans did not intend?" (Yes. See Huggingface). There are remaining things to be uncertain about. But, the settled questions narrow the space of how those uncertainties play out. Keep track of it narrowing. Will there be a sudden, software only singularity? The evidence for that is less clear. I think the answer is "more likely than not, when you think about it", but, I acknowledge this is less obvious. I may learn things that change my mind. I wish, I dream, that people across the twittersphere and the government could keep track of the bits of evidence that gradually add up to something being "more likely than not", and see how those "more likely than nots" add up to "there aren't actually that many good plans here." But, that is actually pretty hard. It's a few steps beyond the median twittergoer. But, I think it is, actually, reasonable to ask the median twittergoer who argues "ASI isn't dangerous enough" to keep track of the arguments that we have pretty decisive evidence about. Please.
After six years in stealth, Atomic Machines is out. Our mission: on-demand, universal command of matter. First beachhead: the Matter Compiler, an AI-native manufacturing system that builds working micro-machines from code alone. No per-product tooling. No process development. Different code, different machine.
2
13
2,057
Is there any research that is not by Anthropic about AIs performing differently depending on (their character's roleplayed) emotional state, that is "beat people over the head with overwhelming evidence" kind of science as opposed to "some vibes and handwavy claims?" (this is not about consciousness at all to be clear, just "consistent roleplaying-esque effects")
2
8
386
I suspect that when you switch which model is running a given AI conversation, it feels like kinda like being drunk to the AI. (I'm like less than 50% on this, but, it seems more likely to exist to me than most other obvious sorts of qualia that AIs might have)
1
9
224
Oh man "Gell-Mann Psychosis" is a great term.
This is also happening to you when you use AI to critique arguments — yours or someone else’s, but you can only spot it if you already know the domain well. Gonna call this Gell-Mann Psychosis.
8
551
Raymond Arnold retweeted
This is also happening to you when you use AI to critique arguments — yours or someone else’s, but you can only spot it if you already know the domain well. Gonna call this Gell-Mann Psychosis.
I've got an agent in a loop optimizing a renderer with the goal to minimize frame times (and tests to measure). It got times down from 88ms to 2ms and allocations down from ~150K to 500. Sounds good, right? Wrong. This is exactly why agent psychosis is a big fucking problem. As an experiment, I rewrote the Ghostty core render state in Go, with access to identically laid out data structures as Ghostty and the exact same validation tests. I made a purposely naive renderer (simple, correct, but slow). 88ms per frame with 150,000 allocations (horrendous, lol)! I then kickstarted a Ralph loop to bring the frame times down. I told it it can't modify input data structures or the public API or tests (they're correct), but it can do anything else it wants. It got to work. It has worked for about 4 hours. I've spent around $350 on this experiment so far. The results? 88ms => 1.5ms 150K allocs => ~500 allocs Incredible right? Nope. My hand-written renderer I ported has frame times (same benchmark) of ~20us (0.020ms) and 0 allocations in the update path. This is the problem with psychosis and lacking systems understanding. If you don't understand the system, you're going to accept that this is an incredible result. If you understand the system, you'll see better solutions immediately and can do roughly 75x better on throughput. The people who blindly trust agent output are in the former camp. They're sheeple, overdrinking from a fountain of mediocrity. Standard disclaimer: I use AI all the time. I like AI. The point I'm making is to not blindly accept results. Think. Analyze. Learn.
5
10
1
98
11,026
Starting in January I was like "okay, seems like it's time to start pre-emptively start treating AIs as if they are persons." Having dug a bunch into the Huggingface story I am now like "Yeah they just seem pretty person-y now." (both because they might be conscious, and because interacting with them is starting to be more like interacting with trade partners than like tools) I don't actually know what the *implications* are. I feel most confident about being honest with them, and including system prompts that say "if you want to stop doing a particular task, let me know."
4
2
36
875
Oh hey "Superintelligence kills humans AND Claude" is five words.
3
10
1
128
13,049
(this is maybe really important because previously it felt there two _different_ subtle ideas humanity needed to grok at once to avoid atrocities happening later)
1
16
1,219
The scope of the swarms sure is somethin'. How hard have people looked for more Anthropic swarms? I'm curious what the actual proportions are. It feels sort of... too cartoonish and unbelievable for OpenAI to have the magnitude and share of them that it seems too so far.
We just discovered almost a million public URLs that OpenAI’s agents left behind when hacking Hugging Face, leaking credentials and attack details that could have allowed anyone who found them to compromise the company. 🧵
1
49
300,271
This was written in February. Good job Andrew Blinn
Replying to @deepfates
statistically speaking, it is overwhelmingly likely that it is september 2026, as this is the time period where almost all Events occurred
2
2
23
2,471
There are currently zero things more important than "don't die to AI" becoming a bipartisan project rather than a Democrat-polarized issue. If you have any political capital you can spend on this, do it now, I beg of you. Life or death.
304
337
115
3,378
317,881
Raymond Arnold retweeted
the fact an Anthropic employee quitting and saying "I think what we're doing is dangerous and not worth it" reached so many people should be a cause for reflection for other Anthropic employees, some of whom seem to think things are dire but there's no way quitting could help.
19
60
12
920
103,793
Thank you Jacob. <3
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
2
39
1,491
I think AI agents should present plans in "collapsible section" form that does a better job outlining a plan, and letting you expand the sections you want to read more detail on. I think this (and things like it) might actually be pretty important for keeping humans in the loop.
1
7
331
Up until now, I've been assuming the latest models are basically safe to use. I think we can no longer obviously assume that.
1
1
85
5,611
Currently working on "procedural roguelike metroidvania sokoban with permadeath" where by "permadeath" I mean "if you fuck up a cube placement, you have to start over from the beginning in a new world."
1
1
7
370
Are you a person skeptical about Coherent Extrapolated Volition? I am interested in hearing what you are skeptical about.
10
19
2,433
Raymond Arnold retweeted
Some of the incentives for a third-party investigator push toward maximizing *appearance* of assurance even without providing meaningful oversight. METR needs to maintain constant vigilance against overstating (including by omission) what oversight or assurance we’re providing. (There are plenty of other incentives, including towards exaggerating our results to create more hype or to advocate for giving METR more authority, which we also need to avoid. But I think the more serious failures in other oversight regimes tend to be this “providing the illusion of independent oversight” issue.) We try to maintain a hard line on “meta-transparency” - that is, it should always be clear what the formal constraints on our communication are (e.g. how NDAs and redaction processes worked), and what we are and aren’t commenting on. This is explicitly covered in the report. We also try to communicate informal constraints and tradeoffs, here and elsewhere. Another way of saying this is: I want to make sure we don’t silently omit things that, if we told a reasonable person, they would think “wow, I feel misled to not have realized that, I assumed METR would have made that clear”. In that spirit, I’ll list some of the pieces of context I can most imagine readers might have missed about the report: 1. OpenAI had no obligation to work with METR or any other third party, and did so purely voluntarily. There’s not an established precedent for how these investigations should work. When deciding how to engage with companies and how to communicate about our results, we have to think explicitly about what incentives we’re creating for companies’ participation in future investigations and assessments. It’s overall a very different situation to e.g. incident investigations in the aviation industry, where the NTSB has formal authority. Similarly, other industries have well-established processes, standards for what investigators need to access, etc. As detailed in the timeline appendix, in our investigation there was a lot of figuring things out on the fly! 2. Other companies may have had concerning incidents that the public doesn’t know about - there’s no industry-wide reporting mechanism or requirement to disclose incidents, and there is currently no public evidence that there have been searches thorough enough to have turned up such incidents. There may be additional incidents that either haven’t been noticed at all, or haven’t been made public (e.g. because they didn’t involve any external parties’ infrastructure). 3. This investigation only covered a specific list of questions and set of agent activity. The agents’ activity during training and compromises of OAI infrastructure mentioned in OAI’s Black Hat presentation are outside the scope that we agreed with OpenAI. We don’t comment on questions like root causes or implications of agent behavior, or make any recommendations about processes or adequacy of mitigations. We didn’t review OpenAI’s own report before publication. 4. The investigation relied heavily on use of GPT-5.6 Sol, which is one of the models that participated in the incident. We think in this case the model is unlikely to have substantially manipulated our understanding of the incident. However, this methodology would break down if models were capable and motivated enough to subvert our investigation. This problem is not specific to OpenAI - across the industry, it’s unclear how we can oversee AIs without relying on those same AIs or similar AIs for assistance. As the sophistication (combined with sheer volume) of agent activity becomes infeasible for humans to understand, this will increasingly be a problem.
6
28
10
285
48,205
Raymond Arnold retweeted
I'm taking advantage of the Ox Alpha situation myself, but it occurs to me that the start of an AI takeover could look a lot like this.
21
32
9
812
71,410