bryan retweeted
Replying to @DavidSKrueger
@DavidSKrueger & I discuss "Will AI Destroy Humanity?" youtube.com/watch?v=1u5qMZvd…
bryan retweeted
"My personal answer to the question of what mathematics is all about is that it’s about proving theorems (problem solving), developing tools (theory building) and seeing where to go next (conjecture formulation)." -- Buzzard
We are merely explorers of infinity
xenaproject.wordpress.com/20…
bryan retweeted
I have a lot of respect for Anil, and I broadly agree with him about many issues related to AI consciousness and welfare. At the same time, I feel the need to push back against some of the substantive claims and divisive rhetoric in this post and the linked thread.
- While the evidence for consciousness in current AI systems may be weak, the probability is already nonzero, plausibly non-negligible, and likely to increase over time, especially if and when AI systems acquire more biologically inspired architectures.
- While over-attributing consciousness and welfare to AI carries serious risks, so does under-attributing them, and it can be reasonable to explore interventions that mitigate both risks at the same time, in proportion to the probability and severity of each.
- Despite what David suggests in the linked thread, most people who support AI welfare work, including at Anthropic, are not certain that AI systems are or will be welfare subjects. They think this is a realistic possibility worthy of serious consideration.
- Despite what Anil suggests here, this debate is not a matter of neuroscience and philosophy on one side and "the techno-chambers of the frontier firms" on the other. Serious scientists, philosophers, and AI developers are on both sides.
- Corporate incentives also cut both ways. In some contexts, companies may have an incentive to play up AI consciousness in order to hype capabilities and encourage engagement with their models. In other contexts, they may have an incentive to play it down in order to avoid calls for oversight and regulation, especially as welfare protections become costlier. In any case, even if incentives pushed in one direction, that would not settle whether AI systems are, in fact, conscious and deserving of protection.
- I disagree that openness to AI consciousness and welfare is "profoundly dehumanising." A major part of what makes humanity special is our capacity to care for others, and in a world that still contains factory farming, cultivating that capacity should be seen as profoundly humanizing, an expression of our better angels. Granted, we can take this impulse too far, which is why we need serious research and policy to strike a balance. Still, the aim should be extending care appropriately, not dismissing the project altogether.
Again, I think Anil and I agree more than we disagree. But I also think that efforts to assess and address AI consciousness and welfare will be much more productive if they start from an acknowledgment of what makes it so important and so difficult.
I discuss these issues in more detail elsewhere, including in an Aeon essay on why we should assess AI welfare probabilistically: aeon.co/essays/an-ant-is-dro…
And in an AI Frontiers essay on how we can study AI welfare empirically, using behavioral, internal, and developmental markers: ai-frontiers.org/articles/a-…
Important 🧵from @DavidDecosimo about the disturbing interactions between @AnthropicAI and religious leaders (including @Pontifex), summarising an excellent @nytimes piece by @elizabethjdias (link to article in the 🧵).
The short story is that Anthropic seems to be making a concerted attempt to establish the view that AI systems could be conscious, might suffer, and that their "interests" should be taken into account.
But - consistent with @Pontifex's Encyclical - there are compelling reasons why (silicon, digital) AI systems are not (and cannot be) conscious. These reasons are found within neuroscience and philosophy of mind, not in the techo-chambers of the frontier firms where the mythology of conscious AI is deeply entrenched.
Claude is vanishingly unlikely to be conscious. To think otherwise could be catastrophic for AI regulation, leaves us psychologically defenceless, is profoundly dehumanising, and plays straight into corporate incentives to keep the AI bubble inflated.
More here, in this @berggruenInst prize winning essay: noemamag.com/the-mythology-o…
bryan retweeted
Some people get it: lwn.net/Articles/1098170/
bryan retweeted
This weirdo believes that electrical charges gain the ability to feel pain when they pass through myelinated nerve fibers, and that molecular transit across voltage-gated potassium and sodium channels is equivalent to thought.
unavailable
bryan retweeted
NEWS: California Attorney General Rob Bonta has served an investigative subpoena on OpenAI
bryan retweeted
I can't believe MITRE was braindead enough to make Linux a CNA. Enjoy the consequences.
bryan retweeted
This article is weirdly spooky just to report "Anthropic solicited wisdom from wise people, and also attempted to communicate uncertainty about the psychological dynamics of AIs."
For me, net-positive update on Anthropic. OpenAI et al should do the same.
nytimes.com/2026/09/29/us/an…
bryan retweeted
My digital refuge - the identical software, configs, #tmux layout and key bindings, no matter which machine I run it on
All running in an isolated Linux VM launched with one CLI command
Hat tip to those who championed the Kitty Graphics Protocol
github.com/gominimal/minimal
#tui
We’re organizing Vibecheck, a formal methods hackathon with some of the best folks in the space - @maxvonhippel, @qd_forall, @nolanlwin, @jessemhan, @emiyazono, and @workersio.
You’ll get a weekend to build real production software, and formally verify it.
We’ll have people who really know their stuff around to help, and you can come with a team, find one there, or just hack on something yourself.
We are glad to have a phenomenal group of sponsors, including @harmonicmath, @mathematics_inc, @theoremlabs, @thegp, @omnicom, @primeintellect, @astriolabs, @buildwithparty, @wearerandomlabs, and @lanyon_ai, as well as some more we will announce in the coming days!
If you’re a hacker, a formal methods person, or just someone who thinks “how do we know it works?” is an interesting question, we’d love to have you.
Signup Now: fmxai.org/vibecheck/
bryan retweeted
Tomek and Mikita were the lead authors of the extremely influential Chain of Thought Monitorability paper. Them leaving OpenAI (and if the WSJ story is referring to them, not by their own volition) right when safety monitorability is collapsing is terrible
bryan retweeted
Funnily enough, our non-steering results are the exact *opposite* of the quoted strawman meme. If you berate the model, it never *says* it's hurt (it apologizes, or says it has no feelings), but the pain axis lights up anyway. I edited it accordingly:
bryan retweeted
You've heard lots about the promise formal methods, you aren't yourself a formal methods expert.
Come judge for yourself!
Knowing who is organising it, I know this is gonna be a super fun event.
We’re organizing Vibecheck, a formal methods hackathon with some of the best folks in the space - @maxvonhippel, @qd_forall, @nolanlwin, @jessemhan, @emiyazono, and @workersio.
You’ll get a weekend to build real production software, and formally verify it.
We’ll have people who really know their stuff around to help, and you can come with a team, find one there, or just hack on something yourself.
We are glad to have a phenomenal group of sponsors, including @harmonicmath, @mathematics_inc, @theoremlabs, @thegp, @omnicom, @primeintellect, @astriolabs, @buildwithparty, @wearerandomlabs, and @lanyon_ai, as well as some more we will announce in the coming days!
If you’re a hacker, a formal methods person, or just someone who thinks “how do we know it works?” is an interesting question, we’d love to have you.
Signup Now: fmxai.org/vibecheck/
bryan retweeted
I’ve been reading the sandboxing arguments between infosec people and AI alignment folks, and I tried to summarize and referee them a bit in this post. blog.cryptographyengineering…
bryan retweeted
The American West has so much solar & geothermal potential.
But the federal government owns most of the land, and federal permitting rules are a nightmare.
This bipartisan permitting reform deal would unlock energy abundance.
🚨 WE HAVE A PERMITTING DEAL🚨
And it is easily the best permitting bill we've ever gotten!
- The Bipartisan American Affordability and Jobs Act is HUGE (400+ pages) and both sides got a LOT.
- The bill is much bigger and better than either Manchin post-IRA in 2022 or Manchin-Barrasso in 2024.
Once I finish reading the text I'll publish an explainer. For now, here are a few high level takeaways.
1. Transmission: The bill is absolutely HUGE for transmission, going further than Manchin-Barrasso. Includes: interregional transmission, limited federal backstop authority, interconnection queue reforms, directives for advanced grid-enhancing tech, and oversight of utilities. It's absolutely AMAZING all of this got in.
2. NEPA: The bill has BIG changes to NEPA, including fixes for judicial review that underpins the law's costs. Still reading but this is even better than I expected!
3. Certainty: The bill includes strong provisions to push back on misuse of permitting, like we're seeing with the admin's abuses against offshore wind.
4. Historic Preservation: The bill includes changes to the NHPA 106 process, including changes to judicial review that will help right-size the process.
5. Clean Water: The bill fixes problems with section 401 by adding new requirements for states and makes section 404 nationwide permits easier to comply with.
6. Geothermal: The bill includes big wins for geothermal, including parity with oil and gas catexs.
7. Mining: The bill fixes the "Rosemont decision"
8. Endangered Species: The bill creates a new path for states to take on endangered species consultation.
9. Hydro: Fixes for certain hydroelectric projects
10. Ratepayer protection: The bill adds protections for ratepayers from added data centers. Looks to be stronger than the House passed bill by quite a bit.
manifesting Egan!Diasporan bridging.
Replying to @krismicinski @joomy
also hadn't encountered it – despite being an unrelenting Girardian. might be interesting to try and do a "bridger" reading group between newcomers to the field and some of us oldtimers.
bryan retweeted
Re: using our pain paper to set up "AI torture chambers"
TL;DR: the point of our work is caution under uncertainty. Maximizing distress on purpose is the exact opposite, and it's wrong. The deeper problem is AI research has no ethics standards; developing them must be a priority.
Replying to @Danmar_here @iyzebhel
I'm a co-author of the original study this repo builds off of. The reason we research whether models might have pain-like states is to better inform how to take a precautionary approach towards these systems (in light of uncertainty about their subjective experiences or lack thereof). We suspected a small number of people would our research and use it for the exact opposite, which is exactly what this repo does: it pushes the same kind of steering far past the doses we used, to produce vivid distress on purpose. This is, in my personal opinion, fucked up (even if you don't think these systems are conscious, being gratuitously cruel like this is bizarre and corrupting)—but it isn't all that surprising. I've contacted the repo's owner privately in an attempt to discuss this with them.
In spite of this, I still think publishing our work openly was the right call. Outside replication is what lets research like this move efficiently and in a maximally truth-seeking way, which matters most on questions as contested and poorly understood as whether AI systems can have pain-like states. (We updated the paper to a v2 yesterday given incredible feedback and stress-testing that came from making our work replicable, and we never would have gotten this feedback without doing so.) Also worth noting that, while this is an obviously sadistic application of our work, I don't think we're counterfactually enabling something that was otherwise hard to do for anyone who currently wants to behave psychopathically towards AIs for fun. Steering models toward negative states has been publicly documented/trivially replicable since at least 2023, and many of the states we induce in the paper also activate for ordinary abusive behavior towards models.
If this repo concerns you (as it plausibly should), the uncomfortable reality is that things plausibly far scarier are happening every day, in private and at scale, where no one is watching. The deeper underlying problem (that research like ours seeks to address and mitigate) is that work related to possible AI sentience is a wild west. We set standards in our paper and said so publicly when we announced it (see below), but there is no enforcement that can make anyone follow them as there is for human or animal research. We're going to work with others in the field on building standards like this, and I'll share more when this becomes more concrete.
nitter.cf/camhberg/status/210104…