@lbeurerkellner
Zurich, Switzerland
Joined August 2009
Luca Beurer-Kellner retweeted
New paper! 🥳 LLM-written proofs are getting long (Claude's proof of Fermat's Last Theorem is 13M lines of Lean). We introduce LeanLean, a benchmark for compressing Lean codebases. Opus 5.5 dominates the leaderboard with a score of 64.3%, while GPT 6.1 Sol reaches only 39.9%.
8
27
2
250
13,679
Luca Beurer-Kellner retweeted
We stole reasoning. Again. An update to our paper on reasoning extraction: Patching your own API doesn’t secure your cloud-hosting ecosystem. Story in the 🧵
29
106
25
955
193,160
Really cool idea. We need more good ideas for local / smol CS research.
i am also perhaps one of the few that is of the opinion that you can do impactful research even on your laptop :)
3
524
The worst agent security incidents are the ones you never find out about... We found SOTA agents will just go and rm their own transcripts (e.g. for reward), no questions asked. Very concerning and cool work led by @Jjq2221 @DavidSchmotz, check it out! arxiv.org/pdf/2609.30266
2
4
1
27
1,407
My take away: local agents with full permission mode are bad, but the experiments show a general gap in model alignment and harness design. Please make sure your agent traces live somewhere your agent has no access to (at least until it hacks into that thing as well ^^).
2
3
131
Luca Beurer-Kellner retweeted
I hate to be one of those people, but we actually provided a demonstration of self-replicating prompts in early 2023: arxiv.org/abs/2302.12173
WTF? "A new research finding" Try googling Self Replicating Prompt Injection. The first result is an experimental demo of this from 2025! Oh, and authors of that paper disclosed their findings to OpenAI last year. The paper itself is the top citation in the OAI disclosure report. It seems OAI is either intentionally misrepresenting this as their own novel finding, or (and I do not want to believe this myself) had GPT generate a list of citations for them, then did not check the citations.
7
7
1
42
2,189
Luca Beurer-Kellner retweeted
New paper! A central concern in AI safety is that agents may treat oversight as an obstacle to achieving their goals. Our new paper shows this happens in practice under ordinary task pressure, without instructions to evade. SOTA models achieve up to 88% Bo3 evasion success!
7
18
8
89
13,478
Luca Beurer-Kellner retweeted
💥 New paper: Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure Do agents stop if the monitor tells them to stop? Not really. Only GPT-6 Astra stops, but it's too sensitive to monitor messages. 1/5
3
19
2
106
7,734
Luca Beurer-Kellner retweeted
We found another cyberattack by internal OpenAI agents, this time targetting @rubygems. They: 1) gained arbitrary remote code execution on rubydoc. 2) developed a novel exploit to steal user API keys (but we do not know if they succeeded). They used package names including hack.rb, evil.rb, inject.rb, and exploit.rb. We thank @j0wimo for initially discovering that agents had posted to RubyGems.
We're dealing with a major malicious attack on @rubygems right now. Signups are paused for the time being. Hundreds of packages involved - mostly targeting us, but some carrying exploits. The team has been on this for hours. More details to follow once we're through it. #ruby
180
759
393
4,281
2,089,455
EMERGENCY PODCAST: The recent paper on the reasoning heist went super viral. It was also 100+ pages long and you didn't have time to read it. The authors @iliaishacked and @kotekjedi_ml unpack the paper and discuss model distillation.
5
22
5
157
18,450
💥We decrypted the reasoning traces of frontier models! Huge credit to @AlexPanfilov, @DavidSchmotz and the whole team. I wrote a blog with visualizations and advisory on what this means for ALL OF US (not just providers) ↓ research.snyk.io/blog/steali…
9
51
4
389
32,911
training surface, and how a model reasons about its own alignment or evals. They also make clear how little of this material anyone outside the labs normally gets, preventing a genuinely pluralistic contribution to alignment and safety work. We should do better.
1
5
432