Nora Ammann retweeted
Some new misalignment disclosures from OpenAI:
• Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further)
• In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks
• A new research finding, demonstrating that one can construct self-replicating prompt injections
alignment.openai.com/misalig…
It's amazing to have @trailofbits on board!
When designing the programme structure, we have a paradox to resolve: how can we go as fast as possible while keeping our work absolutely rigorous? A central red team was the answer we came up with.
No doubt we'll be learning a lot!
We're the only red team for @ARIA_research's Safeguarded AI programme, led by @ammannnora.
Each development cycle, we attack the code, its proofs, and everything those proofs rest on: the models, specs, assumptions, and deployment choices behind every team's assurance case.
Nora Ammann retweeted
I am running a formal methods hackathon on November 1st, with a bunch of my friends. Contestants will compete to build normie software -- video games, AI agents, travel booking tools, etc. -- but formally verified. We are going to measure to what degree these tools are usable by SWEs with no prior FM background and no FM-specific education. This is a super exciting opportunity to see where the ergonomic gaps are that we, the FM community, need to fill in order to take advantage of our moment in the zeitgeist. If you want to get involved, please DM me! In particular, I think I've got location/food/tokens/photography all covered, but I need some sponsors for the prizes. If you make a cool gizmo or product we can give away and you're down to donate one, please do so!
Proud and pleased to be finally able to announce our latest set of Creators working on technically exciting and societally critical projects.
Formal methods have the potential to lead to a phase change in cyber resilience, and it matters that we get there fast!!
AI-enabled cyber attacks are getting faster + going further. But what if AI-enabled formal methods could make high-assurance cyber defence practical at unprecedented speed + scale? We’re funding eight teams with £22m to test that idea. Meet them here: link.aria.org.uk/SAI-TA2-Cre…
I hope the rest of Europe will follow suite. Our climate, and our societies, need nuclear energy.
Oggi il Parlamento ha approvato in via definitiva la legge sul nucleare.
Dopo quasi quarant'anni, l'Italia torna ad avere un quadro di regole per produrre energia nucleare. Una scelta coraggiosa e di buon senso.
Oggi importiamo elettricità prodotta dalle centrali nucleari di nazioni vicine, ma non possiamo produrla noi. Questa legge serve a cambiare le cose.
Cosa significa per famiglie e imprese:
• Più sicurezza energetica: meno dipendenza dall'estero e meno esposizione alle crisi internazionali che abbiamo pagato in bolletta.
• Energia stabile e a basse emissioni, che lavora insieme a tutte le altre, a partire dalle rinnovabili.
• Una risposta concreta alla domanda di elettricità che crescerà con data center e intelligenza artificiale.
Parliamo di tecnologie di nuova generazione, non delle vecchie centrali.
Ora il Governo scriverà le norme attuative su sicurezza, autorizzazioni, gestione dei rifiuti e rapporto con i territori.
Era un impegno preso con gli italiani. Un altro impegno che abbiamo mantenuto 🇮🇹
Nora Ammann retweeted
*Security is sleeping on emerging catastrophic risks*
(cross-post from my blog..)
We've just seen:
* a campaign that used agents to compromise ~100 businesses and steal about 600,000 credit cards with minimal human involvement; token costs were ~$25 per target successfully hacked.
* Hacktron getting access to OpenAI’s monorepo by using Claude to exploit a blind buffer-overflow RCE in a way that (to me) felt superhuman.
* OAI/Huggingface.
* And, of course, there’s the ongoing explosion in newly discovered vulnerabilities.
We've muddled through all manner of crises in security before, and for most AI attacks, cyber attack and defense will reach a natural equilibrium over the next few years, as we figure out how to mitigate cyber attackers who have limited goals like espionage and ransomware.
But some cyber attackers will have maximalist or nihilistic goals, and I'm concerned about this because while yesterday's 'maximalist attackers' (e.g. Russia -> Ukraine, US/Israel -> Iran, Iran -> US) were bottlenecked by labor; today's aren't.
I think it's hard to get true intuition for the shape of the risk here.
As an intuition pump, imagine it’s a year from now -- Q4 2027 -- models are a year better (meaning open weight models are better than today's closed frontier), and in this environment, Iran unleashes a swarm of 100k hacking agents using a safety-stripped (let's say) GLM-5.6.
Imagine the damage such an agent army could do given what we've observed with respect to the paper-thin resistance of today's networks to attacks from today's agents. I suspect the damage from such an attack would far exceed the damage caused by NotPetya ($10 billion USD ten years ago).
Or imagine it’s Q4 2027 and an AI-security PhD student whose name rhymes with Morris, who's researching wormable offensive-agent harnesses in the lab, decides, out of nihilism or sheer recklessness, to release his creation into the wild.
Imagine the size of the resulting, exponentially growing swarm, figuring that these local models, a year from now, will be at the level of today's Sonnet or Opus.
Imagine instead of 800 reward-hacking OpenAI agents we now have 250k worm instances (WannaCry, a 2010s-era worm, had about this many).
The challenge in mitigating expected damages from such scenarios is technical, political, and economic. From a microeconomic perspective, as AI improves, and as we continue not to see extreme catastrophes, we have a growing bubble of unpriced risk in which the security community, CISOs, CEOs, and boards may become lulled into complacency.
We are, of course, already seeing this, as some within security think AI is “just another tool,” doesn’t change the fundamentals, won’t require incredible innovation to rise to the occasion of defending against it, etc. This complacency may fly when thinking about ordinary cybercrime, but it misses the emerging tail risks.
There are three things those who recognize the dangers need to do here:
Catalyze appropriate risk pricing. Try to get organizations informed enough to price this new, fattening and elongating tail of risk into their decision-making. Do this by forming an AI security observatory that distills information about emerging AI risks and broadcasts analyses, damage estimates, and forecasts to decision-makers.
Use regulations and subsidies to ensure that critical infrastructure is paying down the risk. This acknowledges that critical-infrastructure organizations that fail to protect themselves from these new threats can impose the costs of cyber catastrophe on society as a whole.
Develop moonshot technologies that make it cheap to pay down the risk. This acknowledges that the measures we may need to take—for example, rewriting entire codebases using memory-safe languages—may be too costly with today’s technology to reasonably prepare ourselves, and that innovation, some of which may need to be funded by government agencies and some by philanthropic funders like the OpenAI Foundation and Coefficient Giving, is necessary.
The security community isn’t used to thinking in societal-disaster-planning terms and has in many ways become inured to them. But there’s no sane empirical case to be made that the risks aren’t here. I’d love to hear from readers about how you’re thinking about this.
Full/longer version here: joshuasaxe181906.substack.co…
It is high-time to draw on the digital forensics community to help investigate this, and undoubtedly future incidents.
We need the expertise of digital forensics experts to these investigations thoroughly, and we need also them to rapidly adapt their methods to and with AI.
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets.
In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵
Our blog: transluce.org/agent-activity
NYT: nytimes.com/2026/09/23/techn…
Some resources Eoghan Casey from @CKE_Ltd kindly shared with me, as a starting point:
solveit-df.org/
dfrws.org/
Nora Ammann retweeted
I'm excited for this New Directions in Software Technology (NDIST) report on formal methods for security! The formal methods scale-up problem has shifted from technical to market structure (coordinating hardware and open+closed source software makers to spend O($10B)). 🧵
Woooh! We have new Activation Partners 🎉🎉 🎉
They are set to strengthen not just @ARIA_research, but most importantly our Creators, and the UK R&D ecosystem!
Big ideas need more than funding.
Today, five new Activation Partners join ARIA, bringing world-class engineering, AI in Science, and translation capabilities directly into our opportunity spaces – speeding up breakthroughs and helping take discoveries from the lab into the real-world.
Meet Adaptyv Bio, Open Athena, Newlab, Edinburgh Genome Foundry, and BioInnovation Institute: bit.ly/4cSxz2u
Nora Ammann retweeted
Keep an eye out for:
- plan for verifying all software coming very soon to @theoremlabs blog
- verified (toy) sandbox demonstration in the coming weeks
- verified production sandbox on Linux by end-of-year
Theorem co-founders @rajashree + @diagram_chaser on how you actually prove an AI agent can’t escape its sandbox, and what happens when the proof exposes a way out:
Jason Gross: "Anything that the agent does inside the sandbox will not result in some canary file outside the sandbox getting changed."
Rajashree Agrawal: "We intend to have one by the end of the year. We've got a verified sandbox now, and the goal is to keep adding features in collaboration with the AI labs."
Jason Gross: "If the AI can guess the secret root key of the package server, then it can do anything. Either I try to prove that the AI is not going to be able to guess that, or I design the system so that the channel just doesn't allow it to authenticate that way."
"You also want to prove that on most inputs there's no change in behavior. Because otherwise it could make it inescapable by saying, well, sandbox just shuts down as soon as it starts."
@theoremlabs
Nora Ammann retweeted
fuck you, from humanity, to the four selfish twits who filed this lawsuit.
and i mean that from the bottom of my heart.
A new lawsuit claims that Anthropic, OpenAI, SpaceXAI and Google illegally agreed to slow their AI development. apnews.com/article/antitrust…
Nora Ammann retweeted
[Emergency alert]
North Korea has launched a suspected ballistic missile. More updates to follow.
Basic resilience preparations are cheap and worth taking. Food, water, basic medical provisions for a couple of days.
Nora Ammann retweeted
🚨 OBAMA ON AI: "If we are thinking about AI just in terms of how do we cure cancer or get better energy, you can do that without having agentic AI and having it just roaming free in the internet. The reason you are doing that is because you have to market a product that people will pay money for.
That’s a misalignment between what our society needs and the commercial imperatives that these companies are facing, not because necessarily they’re trying to do bad things, but because they’ve got to justify these valuations.
So, that’s one more reason why it is really important for us to have a competent government and a serious bipartisan conversation around this issue, and we have to do it fast. And I would encourage voters to pay attention to this. If somebody does not have a serious plan for how to deal with this, then they’re not meeting the moment, and you should probably look for somebody else."
Nora Ammann retweeted
I'm still shocked that the decision from @huggingface and @ClementDelangue not to pursue legal action against @OpenAI isn't getting more scrutiny.
Especially after they were purchased by @nvidia.
Europe, be ready! This is the reality we're living in; let's NOT try to pretend it away.
🇵🇱 JUST IN: Poland is preparing for war with Putin. “We know he is plotting something”
The most important task now is to prepare the country for various scenarios, Polish Prime Minister Donald Tusk said.
“Poland and the Polish people cannot, under any circumstances, afford to be naïve or engage in wishful thinking. We must base our actions on facts, and those facts are unequivocal,” the politician stressed.
Nora Ammann retweeted
Update from us on the Scaling Trust Arena 🏟️
Excited to be working with @andonlabs, @BT6_Official and amododesign.com*! Check out the post, the draft spec, register your interest, and (please!) give us feedback 🤠
scalingtrust.org.uk/blog/upd…