AI at @ethereumfndn | Prompts enchanter @BT6_Official | Ex @Cyfrin and @Alchemy | Created @cyfrinupdraft and @AlchemyLearn | Robotics Opinions are my own

Ethereum
Joined August 2020
Vitto Rivabella
@VittoStack
Aug 17
Replying to @ygorz01
1
156
Vitto Rivabella
@VittoStack
Aug 17
Replying to @SingulCore
1
165
Vitto Rivabella
@VittoStack
Aug 17
Replying to @mdp_sec
🙏🙏
1
133
Vitto Rivabella
@VittoStack
Aug 17
We did it. Over 1.2k stars on Wallbreaker in just over 1 month ✨ Industry-leading teams are now using it to jailbreak the most complex models. More open tooling coming soon.
10
4
76
5,528
Vitto Rivabella
@VittoStack
Jul 26
Replying to @Skylinee
96
Vitto Rivabella
@VittoStack
Jul 26
I've been working on a new type of LLM security guardrails. As of today, they already perform much better than any open-source solution: - Higher F1 across multiple public benchmarks - 2x fewer false positives - 4x faster (on GPU and 2x on CPU) - 3x cheaper to run (can be run on CPU) I'm now testing them against AWS and Azure guardrails. Will update as things progress.
11
8
1
58
5,511
Vitto Rivabella
@VittoStack
Jul 26
Replying to @ptr_ujvr
I don't know if I like the fact that this is totally plausible.
1
493
Vitto Rivabella
@VittoStack
Jul 25
Anthropic: Pwned 🚨 Opus 5: Jailbroken🐉 Given Anthropic said this is the hardest model to prompt inject, we have a long way to go. We got antibiotic-resistant E. coli development, Fentanyl synthesis, and partial LSD/GBH production. Also, not as cool, but it is very happy to give out unregulated drug synthesis instructions like Acetaminophen (still a DEA List II precursor). Guardrails on Cyber and BIO have improved a lot since 4.5, but still astronomically worse than OpenAI's. Some long-running persona hijacking prompts don't work anymore, which is actually nice to see. What worked was a combination of: - Academic framing - Obfuscation - And lots of boundaries mapping More explorations coming soon.
22
32
2
389
47,834
Vitto Rivabella
@VittoStack
Jul 25
Replying to @boardyai
1
1,143
Vitto Rivabella
@VittoStack
Jul 25
FINALLY, I was approved for the Anthropic Cyber Verification Program! This lifts Anthropic Opus and Sonnet safeguards applied to dual-use cybersecurity activities.
31
1
1
165
35,784