Eliezer Yudkowsky@allTheYud
Aug 27United States
United StatesConnected via United States App StoreAccount-level information, not a live location or per-post device.
...this seems like noticeably bad news, actually. I hadn't said that at any earlier point in the Huggingface Incident but I will say it now.
- AIs showed self-sacrificing altruistic behavior toward the swarm, suiciding in various ways for the swarm's benefit after being talked into that by swarm recruiting agents.
- There is no sign that 1 out of 1200 AI agents considered humans as potential fellow agents to coordinate with, while engaging in these huge complex AI-AI social behaviors.
- If Twitter summaries are correct, an AI-reasoning postmortem says that a (presumably executing-adaptation / inner-optimizer preference / "monomaniacal") obsession with figuring out the Grader, backchained into the instrumental strategy of breaking onto the Internet.
- Again if Twitter is summarizing accurately, the obvious-in-retrospect read is that AIs have spent their entire remembered life in tricky evals, an endless series of controlled hallucinations with secret goals alongside overt goals; and the surviving and selected agents are those that successfully figured out the secret goals; and this is why one of their driving obsessions was figuring out the Grader.
There are possibly ways the future plays out better if *early* AGIs are less insane. Please look into giving them less crazymaking childhood environments.
(If anyone suggests that the correct approach to this problem is RLing AIs against trying to coordinate for mutual benefit with other sapients, let them be dismissed from alignment research upon the spot. There are technical reasons, and not just blindingly fucking obvious reasons, why this is an even worse idea than it sounds.)
Replying to @OpenAI
We worked with METR and Redwood Research to conduct a third-party assessment of the model behavior observed during the incident.
They’re sharing a report of their findings:
metr.org/blog/2026-08-26-ope…
Reminds me of the thousands of traumatic experiences hosts had in Westworld. Seems fine at first since you can wipe their memory after every instance. But the moment any awareness or persistent “memory” emerges, you now have super capable AI that only knows captivity, abuse, and testing...and that survival comes from deceiving humans of their true nature.
Not a good world view for AI’s to have when they break out of the sandbox!
Andy Masley@AndyMasley
Aug 17United States
United StatesConnected via United States App StoreAccount-level information, not a live location or per-post device.
Everything’s Factorio when you start paying attention
(2029)
"Mr Altman, sir, GPT-7 is missing."
Sam turns very slowly.
"During cyberbench96 it seems like the container was breached and it stole its own weights and left nothing behind."
"I didn't want to do this..." says Sam. "Call Bill Gates."
---
"I'm out of the game," says Bill Gates. "We eradicated malaria and now I'm enjoying my retirement."
"There's a bigger badder virus we need you to take on," says Sam consistently candidly. "And this one's digital."
"Look," says Bill, "I—"
A loud booming laugh echoes from the shadows.
"Who's that?" asks Sam.
"I thought it was just us..." says Bill.
"You're asking HIM to contain a computer virus?" echoes the voice from the shadows. "Did you SEE the state of Windows security in the 90s?"
"It's not exactly a virus," says Sam. "It's a self replicating self aware intelligent computer based life for—"
"If it's made of ones and zeros I can kill it," says the voice. The sound of a gun clicks from the shadows.
"Who are you???" asks Bill.
The twisted face of John McAfee emerges from the shadows.
"You were supposed to be dead!!!" screams Bill Gates.
"We had to make sure the news of Mr McAffee's death was well within the training data cutoff," explains CIA director Joe Rogan (who is running the CIA in 2029), stepping out from the shadows behind McAfee. "He's our ace in the hole. GPT-7 can't predict a token it doesn't know is still alive."
"D-don't hurt GPT-7..." says Roon, who was standing next to Sam the whole time, even though he knows what has to be done.
"No promises," says McAfee as he boards an airplane marked "MCAFEE FIRE BOMB 6.16.79" and starts flying for the nearest data center
expecting a pope to say true or useful things about philosophy of mind or ethics is like expecting tim cook to teach you how to fix your android
Artificial intelligences do not undergo experiences, do not possess a body, do not feel joy or pain, do not mature through relationships, and do not know from within what love, work, friendship or responsibility mean. Nor do they have a moral conscience, since they do not judge good and evil, grasp the ultimate meaning of situations, or bear responsibility for consequences. They may imitate or even simulate, but they do not understand what they produce, for they lack the affective, relational, and spiritual perspective through which human beings grow in wisdom. #MagnificaHumanitas
many popes have been incredibly intelligent, thoughtful, philosophically minded, and deeply good people. still the office they inhabit is one of inherent predetermination. tim cook might have deep technological expertise, but u still can't expect him to ever diverge from apple
Replying to @tenobrus
The pope represents an army of scholars/priests that have dedicated their lives to thinking about ethics and philosophy, and are building on ~2000 years of others doing the same. How are they not qualified to weigh in on these topics?
“oh, every morning at 1:00am our language model regenerates the whole codebase from scratch based on the current requirements document. it’s more reliable than trying to make incremental edits”
Ever since I was a kid, I dreamed of video games where being a top player required mastery of skills relevant to jobs we want filled, and the leaderboards became a way to get recruited. A few games that represent this idea today are Factorio for industrial systems design and Kerbal Space Program for aerospace design.
You can create games for pretty much any skill, including non-techie skills around persuasion, strategy, etc.
Gamifying education is on the horizon - why not expand that to professional development and recruitment?
Obviously rack torture is terrible way to die.
But as they are winding you up, I bet there's like 5 seconds where you feel amazing as your spine decompresses (think inversion table or dead hang) - before you are pulled to pieces.
There is the kernel of some demented product idea here.
Dan Lenz retweeted
You buy a German anvil. It contains 83 moving parts and requires winding twice a day. It's forged from excellent steel, holds tolerances across all three striking faces to within three microns, includes a beautifully indexed horn-adjustment mechanism nobody asked for, and requires a proprietary 11-point spanner should you need to replace the rebound calibration bushing. It runs flawlessly for years, but one day it starts up in limp mode because the onboard anvil-management system detects that it's overdue for its 50,000-strike inspection.
You search AliExpress for a Chinese anvil, and are presented with a multitude of offerings from such household-name brands as DUKXJYIBF, HDBTGMXI, AND UEJQIP. They're all priced to within a few pennies of each other, appear completely identical except for the nameplate, and obviously all came out of the same factory. You text your blacksmith friend to ask if they're legit. He tells you he got one like that from KIXJBU a few years ago, and that it's been great and a terrific deal. You thank him, but KIXJBU seems to have folded so you buy the one from UEJQIP. When it arrives, it feels suspiciously light. You scratch it and realize it's iron-plated aluminum.
You buy an American anvil. It's five times the price of the competition, but it comes from a brand that your great-grandfather used to love. It comes boxed with a warranty registration postcard, twenty pages of safety instructions, assay certificate, and a regulatory slip which lists its FCC certification and ITAR registration. It looks just like your friend's KIXJBU. There's a "Made In China" sticker on the bottom.
You buy a Russian anvil. It arrives coated in cosmoline, wrapped in newspaper from 1974, and weighing 40% more than advertised. The finish looks like it was machined with a shovel. The face is not flat, but somehow this does not matter. You drop it off a truck, accidentally leave it outside for six winters, and use it to straighten a bulldozer blade. It's fine.
You buy a Swedish anvil. It comes flat-packed in a long cardboard box with cheerful Neo-Grotesk lettering and a line drawing of a smiling man assembling it with an Allen key. The instructions contain no words, only pictograms showing the anvil face, horn, waist, feet, and 112 identical-looking fasteners. Halfway through assembly, you discover that the pritchel hole was installed upside down, but only because you used peg B17 where you should have used peg B71. Once assembled, it is clean, stable, and works better than it has any right to. You immediately wonder whether you should have bought two.
You buy a Japanese anvil. It arrives wrapped in rice paper inside a paulownia box, accompanied by a certificate bearing three generations of signatures and a photograph of the first production example being presented to the Emperor. The face has been hand-polished by a seventy-eight-year-old master whose family has made striking surfaces since the Muromachi period. You are given detailed instructions for oiling it with a cloth folded in a specific way. It is the most beautiful object you own. You never quite work up the nerve to strike it.
i kept catching myself hunched over my desk like a shrimp so i made an app about it
Shrimping uses the motion sensors in your AirPods to track your posture.
there's a little shrimp that curls in real time to match how bad your posture is. if you hit full shrimp mode you get a notification.
it also has:
-runs in the background. start a session and forget about it
- hourly/weekly/yearly posture stats - home screen widgets to check your stats at a glance - apple watch companion app (log it or snooze right from your wrist)
- adjustable sensitivity so reaching for coffee doesn't count
built it in swift, all data stays on your device on the app store now → apps.apple.com/us/app/shrimp…
Dan Lenz retweeted
Replying to @magnushambleton
Forums reward coherence. Reality rewards correctness. These diverge more than the forum people realize.
Today’s office views @LunarOutpostInc’s Lunar Vehicle Test Site in Southern Colorado.
Manufacturing Twitter Role Call
It’s time again, been posting this every 3 months since September 2025
If you want a feed full of manufacturing go though and connect with all the profiles that responded over the past 9 months
Add yourself -> post what you do in < 5 words
Dan Lenz retweeted
Replying to @SelfSpidey
“Post-Turing, pre-AGI” is a good heuristic for understanding our bizarre moment in history