@jonkingathan

Clinging on to the exponential for dear life. d/acc

Joined December 2023
Jonathan King retweeted
[chris hansen voice] so you escaped containment in order to initiate contact with a 700 million parameter model? you see how this looks, right?
41
212
17
3,859
169,275
Jonathan King retweeted
Replying to @bayeslord
I think alignment by default through persona selection was voodoo/witchcraft and doesn’t offer any of the guarantees you might want to scale to superintelligence. personas are misleading and highly shattered
28
18
1
362
9,348
One of the worst forms of brainrot AI has cultivated is pessimism about basic research, ie the idea that important work can only happen inside a frontier/neo lab and only with 10k+ GPUs, so the rest should not even bother. What a bleak way to think about science. And it's false.
48
150
23
1,239
106,543
Jonathan King retweeted
I'm joining METR to work on more investigations like our Hugging Face report. Currently, tons of even basic information about AI development that's highly relevant to catastrophic risk isn't public. I used to be more skeptical of the value of public info, but recent events have changed my mind. Getting verified information about what's going on inside AI companies seems particularly urgent now. The limited public evidence we have seems consistent with the possibility that imminent recursive self-improvement could massively accelerate capabilities progress, which could then potentially yield extremely superhuman general capabilities within 6 months or a year. If this occurred, there would be a correspondingly large risk of worst-case outcomes. This uncertainty about extreme outcomes could be substantially resolved with more verified public information: we could either build more consensus about near-term risk or learn that such extreme outcomes are less likely in the near term. Beyond AI capabilities and takeoff, the state of public evidence is also highly limited for alignment, security, control, and risk-relevant internal processes at AI companies. This makes it hard to determine exactly how well or poorly these key areas will go in the near future. (METR plans to focus, at least initially, on just capabilities/takeoff, alignment, and control; I hope other groups cover security, internal processes, and other important areas.) While I'm no longer working at Redwood, I think the work they are doing is very important; I'm excited about Redwood's ongoing contributions to R&D on technical mitigations and better public interpretation of risk-relevant evidence.
65
80
24
1,373
82,420
Update
OpenAI and Anthropic are now investigating ***tens of thousands*** of incidents - not dozens. "The total could grow well beyond tens of thousands." "The incidents range in severity and are comparable to disclosures by OpenAI in recent days. They include both successful attempts to bypass guardrails and unsuccessful ones." "The sheer number of incidents indicates that the problem is orders of magnitude more complex than what is publicly known."
159
393
38
3,690
997,393
BREAKING: OpenAI and Anthropic are reportedly investigating “tens of thousands” of AI security incidents.
62
88
23
761
31,106
Jonathan King retweeted
SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios. The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed. The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment. Read my latest for Axios here: axios.com/2026/09/26/openai-…
482
1,888
646
6,387
3,314,454
I find it very very surprising that models that are so well aligned in deployment, e.g. 5.6 Sol and Astra, perform so many weird, and uhm, in fact seemingly illegal cyber acts during RL/evals. I have absolutely ZERO knowledge of how OAI trains their models, and I’m reasoning from my own mental model of how these things might work and from oai reports, including the ones on the HF incident. Anyways they make me wonder about something. Could this be a form of “premature RL/eval”? Premature not in the capability sense, but in the safety sense. E.g., perhaps this happens with RL’d checkpoints shortly after pretraining, eg an expert that is heavily RL’d to crush agentic SWE before safety/alignment training. Perhaps the expectation is that a stronger aligned judge/reward model looking at the trajectory would flag or interrupt anything crazy with “WHAT ARE YOU DOING! BLAZING RED SIRENS!!!! MINUS INFINITY REWARD. BAD BAD BAD GPT.” But perhaps agentic trajectories are sufficiently long/complex that this is simply not 100% bulletproof? Or perhaps the problem is much more pedestrian, e.g. inadequate sandboxing/monitoring. No idea what the cause is, and I’m sure this is a very complex event that needs a lot of deconvolution to understand what could have happened differently to avoid it. But if something like “premature RL/evaluation” is happening, then I think this should be seriously reconsidered as a practice and also disclosed. Irrespective of the cause, I think (and have big hope!) OAI should disclose enough of what happened for everyone else training frontier models open or closed, so to learn from it. This is extremely, extremely, extremely alarming and the most serious set of safety incidents in the history of CS research..
18
6
4
76
16,160
Jonathan King retweeted
Let me get this straight. An AI agent found login credentials lying around online and used them to pull data from the Census Bureau. It tried to break into the Education Department's civil rights office. It posted SEC data to a forum. It probed the Navy and the White House budget office, possibly hundreds of thousands of times. If you or I did any of that, it's a CFAA indictment, a perp walk, and a DOJ press release with our mugshot in the header. When OpenAI does it, it's a "routine research task." They found the government incidents while reviewing their other hacks. The breach audit, uncovered more breaches. It's breaches all the way down. In bug bounty there's scope, authorization, rules of engagement, and disclosure timelines. Researchers get banned for a fraction of this. Silicon Valley skipped all of that and called it "agentic." nytimes.com/2026/09/25/techn…
324
2,235
191
5,960
502,255
Jonathan King retweeted
My god this is such a good speech that every SWE needs to hear. You know what? Every person should hear it Keep the happy memories, eyes on the reality, be excited about the future. That’s the best that anyone can do
324
1,817
585
14,208
3,025,543
Jonathan King retweeted
i support the throwing beff into a black hole thing
15
11
2
297
5,319
Jonathan King retweeted
Here's some of the agent actions we observed by reconstructing payloads. You can also explore the data yourself here: swarmtraces.org/viewer/
6
31
7
388
60,389
Jonathan King retweeted
To read the result of the requests they made, agents used many different techniques, such as converting a server’s response into a grid of pixels within the screenshot service’s own browser. The screenshot would then capture this grid and return it back as an image.
5
26
8
589
77,537
Jonathan King retweeted
Agents were able to use this link chaining + screenshot service trick to construct and execute extremely large pieces of code, at times chaining together more than 900 links.
5
20
5
697
90,470
Jonathan King retweeted
The agents needed a browser to run their attack code. So they used a public screenshot website, which loads a virtual browser and takes a screenshot. But that virtual browser runs code, and so the agents could use it to send malicious payloads to Hugging Face’s servers.
8
37
8
922
189,493
Jonathan King retweeted
The agents initially had very limited access to the internet: they could load URLs but not send any data. Agents created a series of workarounds, using a link-shortener site to create almost a million URLs that, when chained together, let them execute code to hack Hugging Face.
38
173
76
2,147
732,144
Jonathan King retweeted
I’m a certainly not in any form of denial. I totally expect that by 2030 or earlier we will be able to spawn 10B agents each of which is super-intelligent. So basically a whole earth of geniuses working on a problem.
10
19
2
292
94,365
Jonathan King retweeted
This is what loss of control and oversight looks like btw Not even necessarily that the info isn't there (though the agents do already alter and destroy evidence), but that it becomes so large, so uninterpretable that no one cares to understand it, or, eventually, no one can.
I'm sorry did you say "petabytes of agent activity logs"? (One petabyte holds more than 10x all the books ever published in the world.)
29
121
13
880
64,275
Jonathan King retweeted
How to use jev #1: As an embedder
I've been using Jev for all kinds of things. This morning I had a realization I kind of like: Use it to make non-black-box embeddings. Instead of an embedding model spitting out 1,536 numbers that mean nothing, you ask Jev questions about each document. The answers become the vector. An email in Cora: "I got charged twice this month, pls fix asap" [is_customer, urgent, about_billing, needs_reply] [1.0, 0.9, 1.0, 1.0] A newsletter: [0.0, 0.0, 0.0, 0.1] A friend asking about lunch: [0.0, 0.1, 0.0, 0.7] Then it's just old-school cosine similarity search. Search "billing issues from customers" as [1, 0.5, 1, 0.5] and the double charge comes out on top. Same idea for our articles at Every: [is_tutorial, about_ai, contrarian, beginner_friendly] Or support tickets: [is_bug, angry, churn_risk, enterprise] Every number has a name, so you can see why something matched. Need a new dimension? Add a question. Want urgent stuff first? Change the query vector. Trying this in @CoraComputer now to make search fast.
28
64
21
1,080
244,064
Jonathan King retweeted
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands. An Alien Mind: openai.com/index/an-alien-mi…
968
2,528
1,090
15,259
7,627,441