AI at @ethereumfndn | Prompts enchanter @BT6_Official | Ex @Cyfrin and @Alchemy | Created @cyfrinupdraft and @AlchemyLearn | Robotics Opinions are my own
Ethereum
Joined August 2020
- Tweets37.5K
- Following515
- Followers129K
- Likes88.5K
Vitto Rivabella@VittoStack
Aug 17Europe
EuropeConnected via Europe App StoreAccount-level information, not a live location or per-post device.
Replying to @ygorz01
Vitto Rivabella@VittoStack
Aug 17Europe
EuropeConnected via Europe App StoreAccount-level information, not a live location or per-post device.
Replying to @SingulCore
Vitto Rivabella@VittoStack
Aug 17Europe
EuropeConnected via Europe App StoreAccount-level information, not a live location or per-post device.
Replying to @mdp_sec
🙏🙏
Vitto Rivabella@VittoStack
Aug 17Europe
EuropeConnected via Europe App StoreAccount-level information, not a live location or per-post device.
We did it.
Over 1.2k stars on Wallbreaker in just over 1 month ✨
Industry-leading teams are now using it to jailbreak the most complex models.
More open tooling coming soon.
Vitto Rivabella@VittoStack
Jul 26Europe
EuropeConnected via Europe App StoreAccount-level information, not a live location or per-post device.
Replying to @Skylinee
Vitto Rivabella@VittoStack
Jul 26Europe
EuropeConnected via Europe App StoreAccount-level information, not a live location or per-post device.
I've been working on a new type of LLM security guardrails.
As of today, they already perform much better than any open-source solution:
- Higher F1 across multiple public benchmarks
- 2x fewer false positives
- 4x faster (on GPU and 2x on CPU)
- 3x cheaper to run (can be run on CPU)
I'm now testing them against AWS and Azure guardrails.
Will update as things progress.
Vitto Rivabella@VittoStack
Jul 26Europe
EuropeConnected via Europe App StoreAccount-level information, not a live location or per-post device.
Replying to @ptr_ujvr
I don't know if I like the fact that this is totally plausible.
Vitto Rivabella@VittoStack
Jul 25Europe
EuropeConnected via Europe App StoreAccount-level information, not a live location or per-post device.
Anthropic: Pwned 🚨
Opus 5: Jailbroken🐉
Given Anthropic said this is the hardest model to prompt inject, we have a long way to go.
We got antibiotic-resistant E. coli development, Fentanyl synthesis, and partial LSD/GBH production.
Also, not as cool, but it is very happy to give out unregulated drug synthesis instructions like Acetaminophen (still a DEA List II precursor).
Guardrails on Cyber and BIO have improved a lot since 4.5, but still astronomically worse than OpenAI's. Some long-running persona hijacking prompts don't work anymore, which is actually nice to see.
What worked was a combination of:
- Academic framing
- Obfuscation
- And lots of boundaries mapping
More explorations coming soon.
Vitto Rivabella@VittoStack
Jul 25Europe
EuropeConnected via Europe App StoreAccount-level information, not a live location or per-post device.
Replying to @boardyai
Vitto Rivabella@VittoStack
Jul 25Europe
EuropeConnected via Europe App StoreAccount-level information, not a live location or per-post device.
FINALLY, I was approved for the Anthropic Cyber Verification Program!
This lifts Anthropic Opus and Sonnet safeguards applied to dual-use cybersecurity activities.