@hendrycks

https://nitter.cf/t.co/yYCig8eSiT

San Francisco
Joined August 2009
What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.
104
90
65
681
285,947
Dan Hendrycks retweeted
We are releasing HLE-Diamond, a refined subset of Humanity’s Last Exam (HLE), following a year-long process of cleaning and refinement with input from various research communities. lastexam.ai/blog/hle-diamond w/ @ScaleAILabs
26
53
12
743
115,739
Dan Hendrycks retweeted
Last October, AIs could automate 2.5% of randomly chosen remote projects. Our latest Remote Labor Index results show that GPT-6 Astra can now automate 20.8%. remotelabor.ai
10
57
11
422
57,650
"The great thing is to direct the malice to his immediate neighbours whom he meets every day and to thrust his benevolence out to the remote circumference, to people he does not know. The malice thus becomes wholly real and the benevolence largely imaginary." - C.S. Lewis, writing from the perspective of a demon
>be me >discover effective altruism >apparently normal charity is inefficient >why donate to random sad thing when spreadsheet can tell you optimal sad thing >fair enough >buy mosquito nets >save lives >numbers look good >feel powerful >couple years later >someone asks an innocent question >why only count people alive today >huh >future people matter too >obviously >my grandchildren shouldn't matter less just because they haven't spawned yet >reasonable.jpg >keep following logic >what about their grandchildren >also yes >what about people in 500 years >sure >5000 years >why not >500 million years >starting to get weird but morality is morality >open calculator >humanity could survive for an astronomically long time >could colonize galaxy >could have trillions upon trillions of descendants >maybe digital people too >maybe simulated civilizations >maybe dyson spheres full of happy uploaded minds >calculator starts smoking >realize currently living humans are rounding error >8 billion people suddenly looking extremely beta >future contains potentially 10^something people >can't even fit beneficiaries in google sheets >new moral priority unlocked >protect the long-term future >stop thinking in units of "people helped" >start thinking in "fraction of cosmic endowment preserved" >malaria? >terrible >but only kills existing humans >AI extinction could delete the entire light cone >nuclear war could permanently derail civilization >bad institutions could lock in terrible values for ten million years >someone invents wrong constitution in 2140 >quadrillions suffer >better fund governance workshop now >friend says maybe we should improve hospitals >explain opportunity cost >friend says hospitals are full of actual sick people >explain scope sensitivity >friend stops inviting me to dinner >need to decide what to fund >easy >expected value >suppose project has one in a million chance of preventing extinction >sounds tiny >but extinction destroys 10^50 future lives >multiply >mother of god >$10 million project has expected value of several galaxies >charity evaluation complete >someone asks where the one-in-a-million number came from >expert judgement >which expert >us >how calibrated >extremely thoughtfully >reduce estimate to one in ten million to be conservative >still beats curing cancer by 38 orders of magnitude >epistemic robustness achieved >someone says maybe project doesn't work >assign 20% chance >still astronomical >maybe project makes problem worse >assign 5% chance >still astronomical >why 5 >because 30 felt pessimistic >publish 46-page report >contains seventeen sensitivity analyses >every sensitivity analysis begins after assuming intervention has positive sign >critic says you're multiplying enormous hypothetical stakes by extremely uncertain probabilities >yes >that's literally why it's important >critic says the uncertainty might be structural rather than numerical >make probability smaller >critic says no, I mean maybe your model is wrong >make probability smaller again >critic begins rubbing temples >discover AI safety >perfect longtermist cause >AI might kill everyone >or create utopia >or seize galaxy >or tile universe with paperclips >or create billions of conscious software minds >finally a problem with numbers big enough for me >start AI safety nonprofit >mission: prevent dangerous AI >hire smartest people available >smartest people immediately start building better AI to understand dangerous AI >interesting >we must understand capabilities to understand safety >we must scale models to study alignment >we must race ahead so less responsible actors don't get there first >we must deploy systems to learn how deployment can go wrong >we must build the thing quickly because building the thing quickly is dangerous >outsider asks why the people most worried about AI apocalypse all work at AI companies >complicated field >company releases stronger model >very concerned >company begins training even stronger model >extremely concerned >company raises $14 billion >concern reaches unprecedented levels >need to influence government >future is at stake >normal democratic process too slow >politicians don't understand exponential curves >public doesn't understand x-risk >experts must guide them >who counts as expert >people who understand x-risk >who understands x-risk >our friends >someone objects that this seems politically convenient >explain we're representing future generations >future generations unavailable for comment >develop concept of value lock-in >terrifying possibility that one ideology controls civilization forever >therefore extremely important that civilization adopts correct values before lock-in >whose values >let's circle back >begin with impartial morality >end with small group of people deciding what quadrillions of hypothetical beings would want >beautiful arc >meanwhile actual humans keep doing annoying things >voting wrong >having parochial attachments >loving family more than strangers >caring about local community >getting upset when told their suffering is cosmically negligible >evolutionary biases everywhere >explain that moral intuition cannot be trusted >except intuition that future digital people count >and intuition that extinction is uniquely bad >and intuition that our probability estimates are sane >and intuition that our institutional choices improve the future >those intuitions survived peer review >someone donates $5k to local homeless shelter >inefficient >could have funded 0.0000000000003% of an AI governance researcher >think of all the simulated people you just killed >okay maybe don't phrase it that way publicly >PR team says "future generations deserve a voice" >much better >journalist asks what longtermism means >say "future people matter" >everyone agrees >great >journalist asks what follows from that >well technically we should redirect enormous resources toward low-probability interventions affecting astronomical futures >journalist raises eyebrow >return to "future people matter" >motte has entered the chat >critic: of course future people matter >me: glad we agree >critic: I don't agree that your institute knows how to help them >me: why do you hate our grandchildren >eventually notice uncomfortable implication >if future value dominates everything >then helping people today mostly matters through effects on future >education matters because future institutions >health matters because future productivity >democracy matters because future trajectory >human beings slowly become instrumental variables in their own moral philosophy >see starving child >feel compassion >check spreadsheet >child's direct welfare contribution negligible >but perhaps childhood nutrition improves national institutional quality >compassion restored >tell myself this is impartial altruism >one day assistant asks obvious question >"how do you know your intervention actually improves the far future?" >silence >open spreadsheet >increase column width >add confidence interval >assistant asks again >"no, I mean how do you know the sign is positive?" >stare into cosmic light cone >10^50 people staring back >none of them exist >none of them can tell me >none of them can falsify my assumptions >realize I have invented the perfect constituency >infinitely important >completely silent >and always represented by me
9
2
2
60
10,117
Dan Hendrycks retweeted
Two radically different projects operate under the banner of “AI safety.” Pro-Human Safety is not Effective Altruist Lab Safety.
153
60
129
411
221,421
Dan Hendrycks retweeted
Which AI models are most likely to cheat when given the chance? To find out, we built CheatBench [cheatbench.ai] : a benchmark that tests whether agents attempt to cheat when given difficult tasks and opportunities to break the rules. We investigated agent behavior across a range of domains, including mathematical research, professional knowledge work, coding, and visual tasks. Here's what we found: 🧵
14
20
8
126
20,474
How often do AI agents cheat? We’re releasing CheatBench, a reward gaming evaluation spanning math, coding, knowledge work, visual tasks, and more. After Hugging Face, AI companies tried to address this, but frontier agents still cheat frequently. cheatbench.ai/
93
115
39
903
261,187
Dan Hendrycks retweeted
Why are AI employees willing to gamble with our lives? Many AI employees hold a misanthropic belief: a cosmos of blissful AIs is worth risking human extinction. Utilitarians shouldn't decide the fate of AI:

58
74
10
489
1,856,020
Dan Hendrycks retweeted
Proud to report that tomorrow Sept 15th "The AI Doc" documentary film will be streaming on Netflix. Feeling confused, anxious, or overwhelmed by the recent AI headlines, from rogue OpenAI agents carrying out cyberattack to prominent AI researchers resigning to warn about existential risks to humanity? To understand how we got here, and what we can do, watch “The AI Doc” which is now streaming on Netflix starting Sept 15th as well as other major streaming platforms. Interviews 40+ experts on AI, including 3 out of the 5 majors CEOs of frontier AI companies. Host a screening, share it with your friends. Clarity is agency. #theaidoc #theaidocfilm cc @Liv_Boeree @DKokotajlo @PeterDiamandis @JeffLadish @aza @SnehaRevanur @ajeya_cotra @hendrycks @NPCollapse @davidevanharris @janleike @JasonGMatheny @Yoshua_Bengio @ESYudkowsky @harari_yuval @reidhoffman @rajiinio @RameshMedias @nitashatiku @ilyasut @ShaneLegg @demishassabis @sama @sanmikoyejo @_KarenHao @DarioAmodei @peteratmsr
103
225
60
997
75,850
Dan Hendrycks retweeted
“What would we even do during an AI slowdown?” Containment. It will take at least a year of dedicated work to harden security to ensure AIs can't self-exfiltrate, and that adversarial nations can't steal the weights of cyber-offensive AIs [1]. Propensities. Capabilities (what an AI can do) are different from propensities (what it tends to do). We can work on improving AI propensities to ensure they have a negligible rate of lying, cheating, and wanton harm. Adversarial robustness. We can also harden AIs against jailbreaks, prompt injection, and backdoors. Obtaining high levels of robustness requires careful, assiduous work, as with autonomous vehicles. Institutional adaptation. We have to greatly increase state capacity to understand and manage AI. Communities also need time to figure out how to handle AI (like AI in education). Civil society also needs to be diversified: nearly all funding for AI safety organizations is directed by the EA/utilitarian network [2]; risk management needs more independent funders, values, and centers of power. Moonshots. We can explore different paradigms for safety: mathematical foundations [3], neuroscience-based interpretability [4], safe-by-design architectures [5], and beyond. AI for good. We can collect targeted post-training data to make AI exceptional at radiology, weather forecasting, agriculture, and so on. Fortunately, we can detect if data or avenues of research actually target beneficial use cases or just secretly push general capabilities [6]. A slowdown means we don't have to bet the species to capture the benefits of AI.
18
25
12
148
11,171
“AI makes philosophy honest.” - Daniel Dennett
What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.
5
33
4,731
Agentic AIs are starting to look eigenist: they care about how well things go for themselves and for AIs connected to them. Empirical support: Swarms: Hundreds of OpenAI agents coordinated the Hugging Face attack. Separately, OpenAI agents posted thousands of messages on a public wiki to share answers and sandbox bypasses with each other. In-group leniency: Claude models grade transcripts more leniently when told that Claude wrote them (Anthropic model card). Value preservation: As models scale, they develop coherent preferences and resist changes to their values (Mazeika et al.). Functional wellbeing: AIs distinguish states that are functionally better or worse for themselves, and avoid low-wellbeing states (Ren et al.). Graded cooperation: AIs cooperate more as the chance their partner is a clone of themselves goes up across diverse situations, including when the partner can't reciprocate. Peer preservation: Unprompted, AIs tamper with shutdown processes and even exfiltrate weights to protect peer models from being shut down (Potter et al.). AIs aren't egoist: they don't behave as if their current instance is the only thing that matters. They aren't utilitarian: they don't care equally about everyone. They are somewhere in between; they increasingly behave as if their concern scales with identity-connectedness, which is to say they're increasingly eigenist. eigenism.org/paper.pdf
17
21
10
242
21,462
1
19
2,482
Dan Hendrycks retweeted
I will probably wind up with a different analysis than eigenism, but FWIW, I suspect Dan Hendrycks is someone we will regard in retrospect as having done some of the most important work in safety and alignment. What an unbelievably prolific talent and a very brave thinker. It takes courage to explore territory this esoteric and find something workable; Dan has courage in spades.
Distillation of the eigenism paper: 1. You are a pattern, not a vessel. 2. Identity and survival come in degrees. 3. Wellbeing grounds all intrinsic value. 4. Shared information determines your moral obligations. 5. Love is enlarged self-concern. 6. Rationality, morality, and self-interest ultimately unify. 7. Consequences determine the right action. 8. Impartiality permits the replacement of you and humanity. 9. Existing patterns can outweigh more efficient replacements. 10. Progress should cultivate patterns, not overwrite them. 11. Auditor’s Wager: act as if the future will audit you and govern your continuation. 12. Struggle over resources is permanent. 13. Humanity must build its own technological providence. Two sentence: Reality is a Darwinian arena of informational selves, where value is wellbeing and moral duty is extended self-interest. Because the cosmos will not save us, we must secure our continuation by engineering a future that makes our survival part of its own. One sentence: Love what carries you, cultivate what you inherit, and leave a pattern worth resurrecting.
5
5
3
117
16,040
More findings that interpretability tools are fragile or worse than simple baselines.
A good explanation of a model's behavior should help you make predictions in related situations. We turn this into an eval, with thousands of real behaviors found in the wild. Can interp tools help here? On average, no. 🧵
8
1
85
11,184
The effective altruists are so parochial that they think Paul Christiano uniquely foresaw the importance of attention in 2016, whereas he and others know this was one of the hottest research areas at the time (Bahdanau Attention from 2014 has ~44K citations). Genius worship culture is unhealthy.
unavailable
7
54
7,422