@blainedillii
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
This. I would just note that liability alone can’t handle the tail risks Bessent mentions, as labs might be bankrupt before they can be held responsible for the full harm they caused. Insurance mandate helps address this. It also gets us embedded inspectors who are hired by insurers, rather than just METR or other anthropic friends
If only there was a way to get auditors with good incentives…
ai-frontiers.org/articles/do…
And they said an actuarial model wasn’t something to get excited about
There will be watch parties all around the world on release day
The Artificial Intelligence Underwriting Company has raised $55M to audit and insure frontier AI models.
Insurance produces audits the public can trust and understand. The mechanism is simple: insurers need audits to price the risk. They pick auditors they believe will help them price the risk most accurately. Insurers then turn the audits into a price signal that cuts through political debates. Insurers are aligned with the public: if they get the audits wrong, they go out of business. That’s different from today, where labs pick their own auditors, creating public distrust.
Insurance only works if we scale up auditing, and we are building the infrastructure to do that. Millions of agents are being deployed; they leave terabytes of traces and will create thousands of incidents across dozens of risk categories. To scale auditing, we need reliable agents to surface evidence traces for human review.
History guides us. Take cars: auto insurers funded crash tests to price risks; crash tests led us to airbags, lowering the risk. Take electricity: home insurers funded lightbulb tests to price fire risks; today, every lightbulb carries a safety mark. Nuclear safety works similarly.
We’ve been learning by doing.. In the last 6 months, we’ve audited and insured frontier agent builders like Cursor, ElevenLabs, Lovable and Harvey. For ElevenLabs to get the world’s first agent insurance coverage, a Lloyd’s of London insurer required quarterly technical audits by an auditor of the insurer’s choice. We’ve built infrastructure for quarterly technical audits of any agent, working alongside firms like Gray Swan (technical evaluations) and Schellman (audit).
In the coming weeks, we’ll share our vision for how an insurance ecosystem can work for the frontier, and open source the world’s first actuarial model for frontier AI catastrophic risk.
We’re hiring across all roles in San Francisco. Come do your life’s work with us.
Rune & Rajiv
What if instead of trying to regulate AI the same way we regulate viagra, we recognize the need for a thoughtful governance regime that draws precedent from the nuclear, cyber, aviation, and satellite industries on how to mitigate risks in highly technical domains ai-frontiers.org/articles/do…
Ah yes, an “FDA for AI”. Because we can just run randomized controlled trials on agent swarms
What could go wrong when evaluating Highly Persistent Internal Model for Safety and Efficacy
It seems to me like the ideal model for AI governance would be embedded auditors from the government, perhaps from a combination of CAISI, the national labs, nat sec world, etc. who don't have any direct regulatory or other coercive authority, but who can flag things to the executive for emergency concerns and who generally have good awareness of what is going on. And then perhaps we have some kind of well defined procedure with careful judicial review to get temporary restraining orders or preliminary injunctions for imminent harms, medium term injunctions for fairly pressing harms but where we don't want to burden the companies for a super extended period of time (The FRONTIER Act does a version of this). And then all of this exists alongside the insurance infrastructure, and the insurers can use the data that comes from the national labs or CAISI or the Natsec world as they see fit in the same way that the Nuclear Regulatory Commission does a bunch of in-depth inspections, and a lot of that information is used by nuclear insurers to price risk.
Another helpful inspiration from Nuclear world could be INPO, an industry consortium that also does detailed technical inspections that inform insurers. There are spillover benefits to safety research, which means frontier labs might be underinvesting in it because they can’t internalize the positive externalities. Allowing them to form an industry consortium to pool resources toward alignment and security research can help address this. But we shouldn’t make it a self regulatory organization, or give it any kind of binding authority.
Alongside all of this, I’m open to IVOs that audit for 1) risks that are uninsurable because they’re too large of tail risks and 2) risks that aren’t really torts or are poorly handled by ex post liability (surveillance/leverage/persuasion by the labs, as an example). But this may not be worth it, as I don’t want to later on an insane amount of regulation.
I hear lots of AI policy folks say the problem with an insurance mandate is that we don't have time. “Timelines are too short” and while insurance might be theoretically good, in practice the IVOs might just be good enough. I’ve come around to thinking they might be correct. But I guess what I’d say is: let's not let this obscure the fact that the IVO design is deeply imperfect (race to the bottom), and the insurance design is much better, though still not a cure all. So the upshot is that we should implement whatever policy we can get in the immediate term while continuing to work towards good institutional design. That is, we shouldn't let the stopgap of FINRA or IVOs or whatever become sticky. And the fact that it probably will is a huge risk to be aware of
Very much appreciate someone laying out the stopgap plan that we urgently need. Our governance of AI can have stages: I remain skeptical of IVOs’ incentives, and eventually we’ll need something better-designed like a catastrophe insurance mandate for frontier labs. But IVOs could be helpful in the short term before they get captured or fall to bad incentives. It seems like a reasonable order of policies is:
ASAP:
-enabling the frontier labs to coordinate a slowdown by having the Admin promise narrow protection from antitrust scrutiny
-the phone-call-based governance Anton lays out where the Admin tells labs they need to have METR, Redwood, etc. red team their systems
Short term:
-formalize that requirement for third party audits, probably by EO, whatever
-Finra for AI or whatever. Not a particularly great design I don’t think, but better than nothing.
-continued antitrust shield to coordinate a slowdown if needed
“Medium” term (6mos - 1 year):
-ideally and (very doably), get insurance ecosystem up and running to create the right incentives for third party auditors
-if that can’t happen, IVOs are second best
Medium term +:
-we really should design institutions with better incentives than IVOs. Stopgaps are fine, but this technology could be too consequential to leave it up to “I trust METR to do a good job.” METR is great! But organizations change when you give them new jobs and incentives! Gabe Weil’s catastrophe insurance and liability work is the best idea i’ve seen for incentive design, and we will really need that. If for whatever reason (perhaps political) we can’t do that in the immediate future, that’s fine, I’ll take “Finra for AI” or IVOs or whatever. But we shouldn’t let those models become sticky! Eventually, we need those IVOs or whatever they’re called to be hired by insurance companies that actually have a financial motive to get the audits right. So to anyone who says “timelines are too short for insurance”: you’re probably right. That means doing non-insurance policies asap. But it doesn’t change the fact that ultimately insurance is better-designed than IVOs and we should implement a well-designed system with good incentives as soon as we can. We can do X, Y and Z (non of them insurance) to prevent “doom” in the super short term. But if it turns out we don’t all die in the next year, then we can do the right thing and do insurance!
“Long” term (more than a few years away):
-the world might be super changed, insurance won’t be enough, and we’ll need lots of ideas for different institutional designs
-or maybe progress and disruption from AI will be less than expected. Even here, insurance isnt a cure-all, and we’ll need to do more than that. This scenario just means we have more time to do good stuff, and hopefully can do better than haphazard stopgaps
AI progress is outpacing democratic oversight, and the government has started to notice. What now?
My new piece argues that mandated independent oversight and incident investigation should be the immediate response.
writing.antonleicht.me/p/sen…
Some reasons, in short:
(1) Third-party oversight gets at the risks we should care about most: not only the properties of one specific model or deployed product, but the internal norms and safety practices. That's how we get at the internal deployments and dangerous training practices that got us here.
(2) It uses and builds flexible expertise. Third-party organisations are the best concentration of talent outside the labs. They can also recruit and scale faster than agencies. And they're more flexible: no matter what law we eventually pass or who wins the knife fight over the AI portfolio, we can keep drawing on third parties.
(3) It gives government something to do about the next incident. Inevitably, something else will go wrong, and government will want to react. These moments are dangerous and ripe for PATRIOT-Act-style overreach. If the White House needs to send someone into the labs to find out what's going on, I'd want it to be third parties, not the NSA.
(4) We can do this today. It's an election year heading into a split government; legislation will be very hard. Under these imperfect conditions, much of this can be done by White House phone call. We should formalise as we go: through closer government-third-party cooperation, through an SRO, and eventually through legislation. But we can start today.
New piece on how this could work and how to get it right: writing.antonleicht.me/p/sen…
Hearsay, character evidence, unfair prejudice — many rules of evidence are partially motivated by distrust of jurors to properly weigh the strength of different facts. If we could build hyper-rational AI jurors or adjudicators, much of the court system might be worth reimagining
Blaine Dillingham retweeted
Oops. Should've said ~4 months
Blaine Dillingham retweeted
The reason that METR and Redwood are preferred for the Hugging Face investigation is that this incident is much more relevant for its alignment implications than its cybersecurity implications. Everyone gets hacked by external actors; we don't need crash public incident reports for most of those. Not everyone gets hacked by internally deployed models that also autonomously hack external partners.
Accurate decomposition of agent behaviors and analysis of how OAI's deployment configurations and testing practices may have contributed to reward hacking/emergent misalignment are the key objectives of this investigation -- not description of where OAI's DLP, EDR, and reverse proxy systems were poorly engineered.
But yes, METR, Redwood, and Apollo should partner more closely with/hire from Crowdstrike, Mandiant, Deloitte, etc., to upskill in cybersecurity (and vice versa).
The independent review of the OpenAI Hugging Face incident, supposedly a watershed moment in cybersecurity, wasn't done by a cybersecurity firm and the authors have no cybersecurity experience. That's bad.
Here's my new video.
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Agreed that many of the mitigations for insurable risks happen to also mitigate uninsurable risks (like alignment research, good cybersecurity, and overall caution)
I’d just note that insurance can in fact cover many “catastrophic” risks. We could mandate that frontier labs carry insurance for very large amounts, with the government providing a backstop as it does in many other insurance industries. However, instead of insurance covering only the first layer of harm and the government backstopping the tail, insurers and government could insure in parallel: for each dollar of harm up to, say, $1 trillion, insurers pay out 10 cents and the government 90. This caps insurer exposure at $100 billion, but because insurers pay 10 cents on every dollar, the premium they charge labs is 10% of what covering the full $1 trillion would cost. The government charges labs the other 90%—it merely multiplies the insurers’ premiums by nine.
Agreed with @The_DanielKing that the Under Secretary should be able to lower the compute threshold that defines a "frontier model." Additionally, to avoid a race to the bottom, two options: 1) auditors should be randomly assigned rather than chosen by the labs, or 2) we should require the labs to carry catastrophe insurance. The insurers would then have an incentive to conduct quality audits or hire auditors who can.
Replying to @The_DanielKing
@The_DanielKing in Policy Gradients today digging into The FRONTIER Act, the most promising federal framework yet for governing risks from frontier AI.
Call for books and movies: crime families but they're AIs
"agents, with their novel set of characteristics (extreme cyber competency, ability to cheaply read a million words in seconds, persistence), will also probably change the contours of digital crime"
hyperdimensional.co/p/on-the…
Some examples of tasks I would do as a paralegal that seem easy for AI now, but for which the AI tools were still very lacking 1 year ago:
-making 3 or 4 page research memos with basic biographical information and links about judges and opposing parties
-making copies of an image rotated many different ways, each on a different page, for showing a witness during a deposition
-proofreading documents for grammar errors
-no joke: taking concentration values (ppm) from a pdf from a laboratory and plugging them into a spreadsheet to calculate another value
How far AI has come: I was a paralegal, and just 1 year ago, there were no good tools for proofreading briefs. Proofreading. The foundation models couldn’t handle 20+ page motions. The grammarly extension for Word was unusably slow. The bespoke tools I demoed sucked. Now, digesting long documents is trivial for even the free tier of foundation models, and they catch exactly the kind of errors I was scanning for as a paralegal.