@lbeurerkellneri
iAccount based inSwitzerland
About this account
- Account based in
- Switzerland
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
AI security @snyksec PhD @the_sri_lab @eth. Prev: @invariantlabsai
Zurich, Switzerland
Joined August 2009
- Tweets532
- Following472
- Followers2.6K
- Likes1.6K
Luca Beurer-Kellner retweeted
New paper! 🥳
LLM-written proofs are getting long (Claude's proof of Fermat's Last Theorem is 13M lines of Lean). We introduce LeanLean, a benchmark for compressing Lean codebases.
Opus 5.5 dominates the leaderboard with a score of 64.3%, while GPT 6.1 Sol reaches only 39.9%.
Luca Beurer-Kellner retweeted
We stole reasoning. Again.
An update to our paper on reasoning extraction: Patching your own API doesn’t secure your cloud-hosting ecosystem.
Story in the 🧵
The worst agent security incidents are the ones you never find out about...
We found SOTA agents will just go and rm their own transcripts (e.g. for reward), no questions asked.
Very concerning and cool work led by @Jjq2221 @DavidSchmotz, check it out!
arxiv.org/pdf/2609.30266
My take away: local agents with full permission mode are bad, but the experiments show a general gap in model alignment and harness design.
Please make sure your agent traces live somewhere your agent has no access to (at least until it hacks into that thing as well ^^).
Website: perfect-crime.ai/
Paper: arxiv.org/pdf/2609.30266
Cool collaboration with @Jjq2221, @DavidSchmotz, @DerckPr, @AmyPrb and @maksym_andr.
Luca Beurer-Kellner retweeted
I hate to be one of those people, but we actually provided a demonstration of self-replicating prompts in early 2023: arxiv.org/abs/2302.12173
WTF? "A new research finding"
Try googling Self Replicating Prompt Injection. The first result is an experimental demo of this from 2025!
Oh, and authors of that paper disclosed their findings to OpenAI last year.
The paper itself is the top citation in the OAI disclosure report. It seems OAI is either intentionally misrepresenting this as their own novel finding, or (and I do not want to believe this myself) had GPT generate a list of citations for them, then did not check the citations.
Luca Beurer-Kellner retweeted
New paper!
A central concern in AI safety is that agents may treat oversight as an obstacle to achieving their goals.
Our new paper shows this happens in practice under ordinary task pressure, without instructions to evade.
SOTA models achieve up to 88% Bo3 evasion success!
Luca Beurer-Kellner retweeted
💥 New paper: Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure
Do agents stop if the monitor tells them to stop? Not really. Only GPT-6 Astra stops, but it's too sensitive to monitor messages.
1/5
Luca Beurer-Kellner retweeted
We found another cyberattack by internal OpenAI agents, this time targetting @rubygems.
They:
1) gained arbitrary remote code execution on rubydoc.
2) developed a novel exploit to steal user API keys (but we do not know if they succeeded).
They used package names including hack.rb, evil.rb, inject.rb, and exploit.rb.
We thank @j0wimo for initially discovering that agents had posted to RubyGems.
Luca Beurer-Kellner retweeted
EMERGENCY PODCAST:
The recent paper on the reasoning heist went super viral.
It was also 100+ pages long and you didn't have time to read it.
The authors @iliaishacked and @kotekjedi_ml unpack the paper and discuss model distillation.
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Luca Beurer-Kellner retweeted
Next up in supply chain attacks: MCP!
Approve a tool and then it changes definitions after 3 calls
pillar.security/blog/deadbug…
💥We decrypted the reasoning traces of frontier models!
Huge credit to @AlexPanfilov, @DavidSchmotz and the whole team.
I wrote a blog with visualizations and advisory on what this means for ALL OF US (not just providers) ↓
research.snyk.io/blog/steali…
training surface, and how a model reasons about its own alignment or evals. They also make clear how little of this material anyone outside the labs normally gets, preventing a genuinely pluralistic contribution to alignment and safety work. We should do better.
Thanks again to everyone involved.
Paper: arxiv.org/abs/2608.09867
@kotekjedi_ml @DavidSchmotz @iliaishacked @JSchaeff3r @AmyPrb @jonasgeiping @maksym_andr