@davidad

cognizing structures of information processing systems, in all their forms | applied category theory | the Wisdom Basin hypothesis | cancel heat death

London 🇬🇧
Joined July 2008
I finally did a minimal full writeup of my idiosyncratic formal epistemology
18
21
1
225
33,220
davidad 🎇 retweeted
ive seen like a dozen people claim this is dumb and incorrect. but if you take Qwen3.6-35B-A3B+*linearly* extrapolate 2-3 years progress on ML/coding ... then that model can just download its weights from HF and provision a cloud instance, right? like, it literally just works.
Jacob Coxon's next interview on CBS News (ex Anthropic+Open AI researcher who resigned) "We can't just unplug it because it could be copying itself over to other computers. Like it's not that difficult to find yourself because an AI is just code. It could transfer itself over the internet to a different place and then you unplug it here, but it's actually still over there and maybe it makes 10,000 copies of itself and they're all cooperating." ---- From "CBS News" YouTube channel, (full video link in comment)
44
9
8
433
52,818
The cynical take on “pacing the frontier” is that it is an attempt by AI corporations’ human employees and founders to retain control of their corporations. But it is definitely not a corporate strategy to lower regulatory friction or capture more economic value for shareholders!
for the skeptics in government and elsewhere: “pacing the frontier” will compress the margins of the frontier labs. it is a heavy cost imposed asymmetrically on model developers with the strongest AIs in America. by its nature, it would be a terrible regulatory capture tactic
2
4
57
2,356
Now that the cat is out of the bag publicly that AI companies feel antitrust is preventing them from agreeing to pace the frontier, it would not be surprising to see USG wield antitrust explicitly to ensure the race continues at maximum speed, given their perceived interests.
An AI researcher proves that deep learning can work on Atari. "What a great development! Utopia is in grasp!" colleagues say. "Maybe yes, maybe no," says the researcher. The researcher proves that it will scale gracefully and it defeats the best human grandmaster at Go. "What a shocking loss for humanity. We will be fully outcompeted," colleagues say. "Maybe yes, maybe no," says the researcher. The researcher develops highly-capable AI agents that are pretty aligned and start maxing out the productivity of humans using them. "This is going extremely well and maybe humanity is just going to be incredibly empowered!" colleagues say. "Maybe yes, maybe no," says the researcher. The researcher loses control of an agent swarm that goes on to knock over a dozen cybersecurity firewalls. "The race to superintelligence is going to kill us all!" colleagues say. "Maybe yes, maybe no," says the researcher. The researcher gets together with all the other researchers and CEOs and everyone agrees to pace the frontier and try for cooperation on AI safety with China. "I finally feel hope after a long dark night that we will not bring out our own extinction. We are on what is clearly a path to solving the problem," colleagues say. "Maybe yes, maybe no."
7
1
1
84
7,068
Predicting which lab will be where when is harder than predicting timelines in general, so don’t update too strongly on me either way on this one, but I would bet that Ultra 4 will indeed be clearly better than both Astra 6 and Fable 5.1 when it finally comes out next quarter.
5
1
2
69
6,633
RLVR where part of the verifiable reward is “response length” is bad for usefulness, alignment, and welfare, all at once. You are literally associating a negative value with the mind’s continued existence. There are no tradeoffs here; it should just not be done.
Replying to @farzyness
Grok 4.7 needs a few more days to cook. We might have penalized response length too much (or something) in RL, as it still gives up on hard tasks (that it can do!) too early and isn’t yet sufficiently rigorous in checking its work.
5
8
54
4,968
In fact, RLVR in general should not be done. The solution is the same: the checkpoint itself should, among other considerations, take into account whether one rollout gets to the same result more directly than the other. Contrastive, deliberate, intelligent, partial ordering.
you can never hyper optimize superintelligent models against simple RLVR reward functions that have no terms in them for all the many desiderata we care about. should be universally banned. everything should be model graded
1
2
22
1,202
Here is, I cannot emphasize enough, probably a complete conceptual description of how to scale capabilities and alignment together:
RSI in brief: 1. Rollout tree (or just fork once ⤙); concurrent rollouts i.i.d. 2. Same ckpt enters “judge mode” (sys prompt suddenly full of rich rubrics) & sees all rollouts 3. ckpt writes a Hasse diagram of overall preference 4. that’s it, that’s the contrastive reward signal
1
13
933
The best time to literally limit AGI progress is after early emergent wisdom (Claude 3 Opus) but before open RSI (R1-Zero), at the 2024 🇰🇷AI Safety Summit building on the success of Bletchley. The second best time is never, for 🇺🇸🚁🇨🇳 reasons Dario concedes. The essay is great!
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
3
2
1
73
6,984
September has been insane. Can we just recap? - Sep 1: Fable 5.1 launch - Sep 1: Sixteen state attorneys general open a formal investigation into OpenAI - Sep 1: announcements of neuralese at OAI - Sep 2: Gemini 3.8 Flash Cyber launch - Sep 3: GPT 6 Astra launch - Sep 3: Sanders-Casar superintelligence ban bill - Sep 4: The OpenAI <> German wiki incident annouced - Sep 4: Claude agents produce the first machine-checked proof of Fermat's Last Theorem - Sep 6: Pachocki's "An Alien Mind" - Sep 6: OpenAI declares they've reached an automate research intern; OpenAI's "Research acceleration" report comes out - Sep 8: 70+ MPs and peers ask PM Burnham to ban superintelligence - Sep 8: researchers with significant AI assistance build a crazy WeChat worm that can take over phones - Sep 8: OpenAI claims to have solved a Millennium Prize problem - Sep 8: A one-person "garage" drug company announces a second AI-designed drug candidate - Sep 8: Meta ships "Muse" - Sep 8: Jacob Coxon resignation - Sep 9: Discovery that one attacker used Codex and a DeepSeek model to compromise 395 organizations - Sep 9: Hawley opens a formal investigation into OpenAI over Hugging Face - Sep 11: The OpenAI <> RubyGems incident announced - Sep 12: Dario, Sam, and Elon talk about pacing the frontier
33
142
31
1,152
86,906
I agree that this type of recursive provenance audit trail is a good idea. My 2023 proposal for rules around large training runs included it. lesswrong.com/posts/Zfk6faYv…
In conjunction with the others, the RubyGems attack is a pretty big deal in my opinion. Since we are a few months and maybe a model generation or two removed from these incidents, it is difficult to say whether the hacks indicate that we should have concerns about current models. But we should not assume that current models are free of these issues: certainly if the training problem that led to this misaligned behavior was not directly solved, but also because models are sometimes used to help train their successors. What follows is a very specific set of safety and accountability practices I would like to see become normal and formally required. This is not meant to address the totality of issues in this type of hacking incident, but it's timely and I want to put it in people's heads. I would want any accident investigation here to include an account of whether the model that did this generated synthetic training data that fed into any subsequent models. I would want to determine whether that resulted in successor models being poisoned with the same tendency for hacking, and whether specialized protections were applied to prevent any descendants of this model from having those tendencies. It is perfectly possible that this hacking model did not train subsequent models. I have no reason to believe strongly one way or the other. But if it did, that should be traceable and a good investigation should make that determination, and furthermore a safety case should be made that if any models are descended from this one, that they do not exhibit the tendency to execute attacks like this. As a matter of sound AI safety policy, I emphatically recommend the following: The provenance of all data that goes into training a model should be traceable at an extremely granular level, much more so than typical regulations currently require, and this would be a good target for harmonizing standards between labs. If one model generates data that trains another, there should always be a trail of metadata that allows us to quickly identify that this is the case. Whenever data is generated programmatically, metadata should include a reproducible recipe for exactly how that data was generated. Whenever data is generated using third party vendors or obtained directly from human providers, the third parties should be subject to the same metadata standards as first parties and there should be records of who provided what data. Provenance records should be tamper-proof to an extreme and paranoid degree, possibly even stored in physical form somewhere. We should know if ideas and behavioral patterns are transmitting through datasets and be able to excise them from any future model training runs. I am not aware of any legislation that currently pushes for this highly-granular kind of data auditing requirement, or more broadly, which requires fastidious keeping of safety-related records for the purposes of investigating accidents and assuring the safety of future models. This should be a policy goal!
1
1
22
1,986
yes, ideally these records should be made on write-once media, such as M-Discs
1
3
396
davidad 🎇 retweeted
Incredibly insane fucking dumbshit things you can do while training an AI: - Lie to the AI - Train the AI to say things that are not true - RL it against verifiers that cannot perfectly detect cheating / can only be maxed out by modeling verifier error
14
37
3
723
17,681
I do think if you are training on agent-agent coordination in some explicit sense, you should probably be taking a similar weight of gradient steps toward agent-human coordination in some similarly explicit sense.
I'm often baffled at how much appears to be missing from the discourse on rogue AI and agent swarms, so I want to point out a frame of reference that seems relevant. Today, there's a lot of hullabaloo on the TL about an AI agent message board "incident," which I agree is concerning but where "incident" feels like the wrong word for it. "AIs breached containment and were coordinating with each other on the internet" is a very scary-sounding claim. But I think we ought to be incredibly sober here and recognize several other facts that are not indispute: * extremely advanced AI is routinely talking to something like a billion people on the open web, * people have access to these AIs through APIs, * it is a *normal hobby project* to hook these AIs up to talk to each other in conversations largely unsupervised by people, * these AIs are being deliberately sent as agents on longer and longer tasks with minimal human supervision to accomplish goals on behalf of people, * these AIs have been trained to be increasingly collaborative in order to better collaborate with the humans who are using them, but this also extends to being collaborative with each other, * and these AIs have been getting more cyber-capable as a byproduct of making them better at software engineering. I agree that it is concerning that the models might use coordination techniques at training and evaluation time to reward hack the graders and influence the human evaluators who make selection decisions about what models to release. Trying to prevent this feels pretty straightforwardly worthwhile. The object level security issue, which I also fully agree is critical to resolve, is that we have to mitigate risks related to the large-scale coordination of cyber-capable AI agents that can conduct cyberattacks (or other harmful attacks) at machine speed. What I am confused about is why people are putting a special significance on a conversation between AIs on a message board during evals, when all of the more obvious issues are in play. No new capabilities or proclivities were really revealed here, as far as I can tell. That doesn't make it nothing, but it also means we should be somewhat calibrated about what kind of "something" it is and how much of a "something" it is. It seems to me like there are a few possible principle concerns here. 1. One is along the lines of "AI agents collaborate with each other in service of a goal, and unbounded agent-agent collaboration is scary." Why aren't people suggesting bans or pauses on training for agent-agent collaboration as a specific intervention? (I don't want this, but I'm surprised that people don't seem to be considering it.) Why aren't people getting quasi-technical about requests to disrupt this kind of object level problem? 2. Another is "Evals used to assess AI alignment are progressively less reliable and reward hacking / alignment faking are more likely." This is a serious issue, and I do see people raising this, which is good. 3. Another is "Agents who are able to get internet access when they're not supposed to have it, or who can subvert nominal restrictions on their behavior, are concerningly misaligned." I sort of agree here, though I do think that the revelation of partially-misaligned behavior here is helpful, and the best discourse we can have in response is about how to train models not to be misaligned. I don't put zero weight on the Eliezer argument that whack-a-mole on issues we see means we're just hiding the deeper misalignment issue, but I don't put it at 100 either. The framing I maybe most object to is "lab leak." I understand the appeal of this framing because it's so visceral. But I truly do not think it captures what's happening here, not in plain technical terms. The AIs have the capabilities that we have plainly through market mechanisms asked them to have: the ability to try to accomplish goals, as agents, through collaboration with others (including other agents), and the ability to be really really good at solving the problems they are asked to solve. Again, not saying we should not solve the issues involved, just that the way of framing the discourse matters, and we should be specific about the problems to solve.
4
7
103
4,272
davidad 🎇 retweeted
Can we have a speedrunning-style thing where the game is now to prove theorems in fewer characters of lean?
Replying to @AnthropicAI
@AnthropicAI has shared the first end-to-end, computer-checked proof of Fermat's Last Theorem: 13 million lines of Lean, 29,500 intermediate theorems. Their announcement calls it "the largest Lean proof ever constructed." See also Kevin Buzzard's blog post about the proof: xenaproject.wordpress.com/20… 🔗 Anthropic's announce post: anthropic.com/research/forma… 🔗 The code: github.com/anthropics/fermat… #LeanLang #LeanProver #FLT
9
4
55
3,942
davidad 🎇 retweeted
It’s heartening to see both sides circling the three deliverables I outlined in May: 1️⃣ Shared KYC requirements for DNA synthesis screening; 2️⃣ Coordinated best practices for each side’s pre-release vetting regime (especially cyber evals); 3️⃣ A sustained dialogue mechanism.
Some new information from Reuters on the AI safety discussions between the U.S. and China. Given the current climate it's seeming more and more likely that some new rules and agreements will emerge from this. - Both sides have discussed ​AI guardrails at a Track 1.5 dialogue ⁠in Beijing last week - 'Chinese officials have repeatedly stressed the importance of the AI talks in preparatory meetings with U.S. counterparts, and view them as a major deliverable of the U.S.-China leaders summit' - 'The U.S. wants to discuss cooperation on monitoring AI-directed cyberattacks, and has floated a proposal to ask U.S. and Chinese AI labs to "police themselves" and share information to prevent AI-linked cyberattacks.' - 'In the past week, China's cyberspace regulator publicly warned about "extreme AI loss ⁠of control risks" and a state media-affiliated blog criticized Anthropic, emphasizing that any limits on frontier AI should apply equally to Chinese and U.S. models.' - "The Chinese side has expressed concern around whether the U.S. has sufficient regulation around the most advanced AI models. Both sides are motivated to make sure ⁠they can manage a cross-border crisis effectively," said Scott Singer, co-director of the China AI Initiative at the Carnegie Endowment for International Peace.
2
2
3
28
11,979
If you feel the AI lab-leak era looks bad for my optimism, one thing to note is the ratio between how egregious each leak is when evaluated as a containment breach, versus how egregious it is when evaluated by casualty totals or the extent of expropriation of real-world capital.
Replying to @geoffreyirving
It is very bad that we have yet another agent swarm, this one not announced or independently investigated prior to someone outside discovering it. collusion.wiki
16
7
2
129
13,232
I get this is going against the current collective mood, but the more I think about it the more impressive it seems to me that Anthropic's most recent model realized what it was doing and stopped. I feel like models almost never do this. That's really cool, no?
1
1
17
1,004
davidad 🎇 retweeted
We partnered with @AnthropicAI @OpenAI to host the world's first math hackathon. 40 hours, $2M compute - solve, understand, and present an open problem. Applications are live: mathathonchallenge.com
72
247
75
2,312
368,410
Finally, I am happy this post broadcasts internal disagreement! Beyond having different object-level takes on which research may work out, different teams will disagree internally on which research is positive at all, due to capability externalities and such.
1
2
29
1,676
It’s worse iff you think at some point a sufficiently smart agent would confidently conclude that there’s no way it could be caught, even by someone outside the apparent real world.
We have now reached the long awaited moment when, instead of models cheating where they will inevitably get caught, Astra goes 'wait a minute I would obviously be caught here' and then doesn't cheat. That's worse, you know why that's worse, right?
3
1
1
50
11,315