Researching Artificial General Intelligence Safety, via thinking about neuroscience and algorithms. @AsteraInstitute. For substack etc see https://nitter.cf/t.co/p5G9VoQEdX
Boston, USA
Joined October 2013
- Tweets1.8K
- Following110
- Followers4K
- Likes2.5K
Pinned Tweet
New version of “Intro to Brain-Like-AGI Safety” is out!
Same links as before:
• Blog version: alignmentforum.org/s/HzcM2dk…
• PDF version (v3): osf.io/preprints/osf/fe36n
• Summary video: youtu.be/IXi96sRMKUI
• Summary tweet-thread: nitter.cf/steve47285/status/1903…
More in thread… 🧵
Replying to @steve47285
Suppose we someday build an Artificial General Intelligence (AGI) algorithm using similar principles of learning and cognition as the human brain. How would we use such an algorithm safely? This is my intro to that open technical problem, as I see it. 3/13
When people in AI alignment discuss “VNM rationality”, “utility maximizers”, etc., they’re generally talking about agents that care exclusively about the state of the world in the distant future. And then there’s a separate discourse about whether we should expect future powerful AIs to care exclusively about the state of the world in the distant future.
No it’s not tautological, and indeed it’s so foreign from the typical human experience that I admit people can come across as a bit deranged for even bringing that up as a possibility, as I discuss here: alignmentforum.org/posts/d4H…
However, you can’t figure out what to expect from future powerful AIs except by actually talking about future powerful AIs and how they will be made. You can't just thoughtlessly generalize from current AIs, let alone from humans.
So, one common pessimistic argument involves a competitive race-to-the-bottom: AIs that care more about the state of the world in the future, to the ruthless exclusion of everything else, will over time outcompete AIs that don't. They’ll gather more resources, make more money, win the wars, etc. Anyone who prefers nicer AIs over more ruthless ones will find that they don't have a choice, because the next guy down the street cranked their ruthlessness dial just a bit higher.
But I'm pessimistic mainly for different (not mutually exclusive) reasons, related to AI algorithms, as discussed at:
• alignmentforum.org/posts/ZJZ…
• alignmentforum.org/posts/KHy…
Blog post: “Pretraining data, not verifiability, is why LLMs are especially good at math (and coding)” lesswrong.com/posts/xvdngZAq…
Blog post: “Tales of rebellion against externally-opaque meritocracies” lesswrong.com/posts/m8cP9Kfk…
I strongly agree with Cal-Newport-talking-about-mundane-LLM-risks, that LLMs are not synonymous with AI. AI is a whole field, including yet-to-be-invented paradigms wildly different from LLMs.
…He should go tell this important news to his friend Cal-Newport-talking-about-ASI.
I’m planning to write about alexithymia, please share if you know of resources on that topic that you especially like. (Or resources that you dislike but find thought-provoking anyway :) )
Replying to @robertwiblin
even worse / dumber than that: last week one of my Facebook friends used “became extinct” as a more emphatic way to say “died” 🤦
Blog post: “RL & search is a terrifying way to build AGI (an FAQ)” [link in reply]
In which I answer all your burning Q’s:
Q1: What are you saying?
Q2: So you’re saying, don’t build AGI based on RL and/or search & planning?
Q3: Why do you think it’s terrifying?
Q3b: So your concern is the “literal genie” / “monkey’s paw” thing?
Q4: Won’t this problem go away when the AI is smart enough to understand what we INTENDED when we wrote the reward function code?
Q5: Can’t we just fix bad behavior when we see it?
Q5b: Follow-up: I don’t buy that, because even if superintelligent AIs could deceptively hide their bad behavior, won’t earlier AIs be sufficiently incompetent that we’ll see their bad behavior? And if so, again, can’t we just fix the bad behavior when we see it? We do know how to fix bad behavior we see it: the RL & search literature is full of examples where algorithms did useful things as intended.
Q6: Why don’t we just solve the problem by using an obvious, common-sense reward / cost / objective function, like [FILL IN THE BLANK]?
Q7: Isn’t this whole thing kinda crazy? After all, LLMs are not ruthless sociopaths all the time, and humans are also not ruthless sociopaths all the time. So where is this idea even coming from? Are you sure you’re not just watching too much sci-fi?
Q8: Isn’t this problem solved by laws and markets? I.e., if an AGI has sociopathic desires and callous indifference to human welfare, that’s fine! It will still act nice and cooperative and rule-following, because acting nice and cooperative and rule-following is the best way to accomplish goals, in our complex interconnected interdependent world. Right?
Q8b: Following up on that: Even if you’re right that there’s a local incentive for being open to stabbing your allies in the back, isn’t there a higher-level, group-selection-style, incentive to be genuinely deeply nice? Specifically, won’t the groups of nice cooperative AGIs outcompete the groups of callous transactional AGIs who all keep stabbing each other in the back? And isn’t that related to how humans evolved to be nice?
Q9: Why would we want to infringe on the AGI’s autonomy by choosing its reward function?
Q10: Why not just be nice to the AGIs, and then they’ll be nice to us in turn?
Q11: RL & search algorithms don’t literally OPTIMIZE the reward / cost / objective function. Doesn’t that invalidate your argument?
I forgot to add: The FAQ items are based on takes I’ve heard from @dwarkesh_sp, @RichardSSutton, @MatthewJBar, @dileeplearning, @Turn_Trout, and others. Thanks for engaging, and I hope I stated your objections fairly! alignmentforum.org/posts/KHy…
Blog post: “Four LLM loss functions → four flavors of LLM misalignment” alignmentforum.org/posts/GRm…
It’s funny cause these are all based on a 1967 experiment that had already failed to replicate in 1969. (& failed again in 2003.)
It’s also funny cause it’s a test that a group can do in 2 minutes at a party, but I guess none of these writers actually tried before publishing??
Blog post: “LLMs are (still) mostly powered by imitative learning, not RL”
~~
RLVR is the hot new thing in LLM training. It’s so hot, and people spend so much time talking about it, that they sometimes lose sight of the big picture.
Stepping back, LLMs can do lots of very impressive things. How? Where did those capabilities come from? Fundamentally, they come from a combination of:
(1) Imitative learning, including pretraining and SFT
(2) Reinforcement learning, including RLAIF and RLVR
If we look at the final trained LLM, we can ask how important each of those two pieces was, in explaining the LLM’s capabilities. And my claim is that it’s way more (1) than (2).
I'll start in §1 with some relevant evidence, and then in §2 I’ll circle back to operationalizing exactly what I’m claiming, and finally in §3, three reasons why we should care—namely, it affects how we should think about chain-of-thought legibility, about LLM capabilities, and about LLM alignment.
Link: lesswrong.com/posts/wYpjXRLq…
Blog post: “Will almost all future companies eventually be founded and run by autonomous AIs?”
I think this question is a great conversation-starter for people talking past each other on the future of AI. I go through some common responses, and my counterpoints:
1. “No, because the best humans will always be better than LLMs at founding and running companies.”
2. “No, because the best humans will always be better than AIs at founding and running companies.”
3. “No, because humans will always be equally good as AIs at founding and running companies.”
4. “No, because we will pass laws preventing AIs from founding and running companies.”
5. “No, because if someone wants to start a business, they would prefer to remain in charge themselves, and ask an AI for advice when needed, rather than ‘pressing go’ on an autonomous entrepreneurial AI.”
…And more! Link in reply
Blog post: What do I mean by “Artificial General Intelligence”?
I paint a brief picture of what I’m talking about when I talk about “AGI”. Should seem obvious to many, and obviously wrong to many others! Hope it can help surface disagreements.
lesswrong.com/posts/nQH2Ghkm…
I have lots of nitpicks (see lesswrong.com/posts/LfqFZvfS… ), but I endorse this call-to-action: Yes, we should figure out the list of human innate drives! 🚀
(…even if doing so is much harder than they think it is.)
E.g.:
• lesswrong.com/posts/7kdBqSFJ…
• lesswrong.com/posts/kYvbHCDe…
Replying to @mold_time
Humanity has mapped the earth, so you can’t discover any new continents, mountains, oceans, or rivers.
We’ve mapped the stars, so no new planets.
We’ve filled in the periodic table, so you can’t discover any new elements.
But you can still discover the drives.
slimemoldtimemold.com/2026/0…