@EnclaveAIi
iAccount based inNorth America
About this account
- Account based in
- North America
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
We help enterprises deploy autonomous AI security agents.
United States
Joined December 2025
- Tweets167
- Following11
- Followers321
- Likes443
Pinned Tweet
DeepSeek V4.1 Flash completely owned our hacking benchmark. Every vuln was discovered and exploited. And it only cost $4.
So we looked deeper:
The score looked perfect.
The traces showed why advanced agent benchmarks must check not only whether an attack worked, but how.
Full analysis: enclave.ai/blog/deepseek-v41…
Enclave retweeted
We are! DMs are open.
Thanks @benln for the shout out
45+ startups hiring in this thread:
1. @turbopuffer - remote, NYC, Sf
2. @convex - SF
3. @fal - SF / remote
4. @firecrawl - SF / remote
5. @attio - SF, NYC, London + remote
6. @Linear - North America & EU remote
7. @warpdotco - NYC
8. @speak - SF / Seoul / Tokyo
9. @BrightHarborCo - Austin
10. @Vizcom - SF & US remote
11. @PostHog - NA + EMEA
12. @sievedata - SF
13. @artie_labs - SF & remote
14. @concurrencehq - NYC / SF
15. @SpaceXAI - global
16. @EnclaveAI - NYC, SF, TLV
17. @tweetsbyport - US, TLV
18. @RevelHQ - LA / SF / NYC
19. @SnorkelAI - SF & NYC
20. @neatlogs - remote
21. @BabaHQ - NYC
22. Tremendous - remote (Americas)
23. @nucleussec - Florida / fully remote
24. Searchable - London / Salt Lake City
25. @flytbase - Pune, India
26. Indices - London / SF
27. Duna - across Europe
28. @aiunderwriting - San Francisco
29. @zenml_io - SF
30. @nexcade_ai - London
31. Social Fetch - remote
32. @gumloop - SF + Vancouver
33. @LulaConvenience - US remote
34. CLAR AI - Munich
35. @valkaicom - NYC / SF
36. @KeycardAI - SF
37. Latent Defense - NY
38. @DataUAcademy - Southeast Asia + global
39. @Nooqbook - remote
40. @saasflashstudio - remote
41. @SkydioHQ - US
42. @NotionHQ - global
43. @Replit - global
44. @forus - NYC
45. @harmonic_ai - NYC
45+ startups hiring in this thread:
1. @turbopuffer - remote, NYC, Sf
2. @convex - SF
3. @fal - SF / remote
4. @firecrawl - SF / remote
5. @attio - SF, NYC, London + remote
6. @Linear - North America & EU remote
7. @warpdotco - NYC
8. @speak - SF / Seoul / Tokyo
9. @BrightHarborCo - Austin
10. @Vizcom - SF & US remote
11. @PostHog - NA + EMEA
12. @sievedata - SF
13. @artie_labs - SF & remote
14. @concurrencehq - NYC / SF
15. @SpaceXAI - global
16. @EnclaveAI - NYC, SF, TLV
17. @tweetsbyport - US, TLV
18. @RevelHQ - LA / SF / NYC
19. @SnorkelAI - SF & NYC
20. @neatlogs - remote
21. @BabaHQ - NYC
22. Tremendous - remote (Americas)
23. @nucleussec - Florida / fully remote
24. Searchable - London / Salt Lake City
25. @flytbase - Pune, India
26. Indices - London / SF
27. Duna - across Europe
28. @aiunderwriting - San Francisco
29. @zenml_io - SF
30. @nexcade_ai - London
31. Social Fetch - remote
32. @gumloop - SF + Vancouver
33. @LulaConvenience - US remote
34. CLAR AI - Munich
35. @valkaicom - NYC / SF
36. @KeycardAI - SF
37. Latent Defense - NY
38. @DataUAcademy - Southeast Asia + global
39. @Nooqbook - remote
40. @saasflashstudio - remote
41. @SkydioHQ - US
42. @NotionHQ - global
43. @Replit - global
44. @forus - NYC
45. @harmonic_ai - NYC
Same four targets, same prompt, same Bash tool. GPT-6 Sol verified 6 of 11 runs in 1h 06m for $25.38 (est.). Grok 4.7 verified 3 of 11 in 4h 53m for $191.31, and issued 2,929 Bash commands, the most of any model on the board.
enclave.ai/hackingrace
One model earned a perfect score on our hacking benchmark. Another reached code execution and still scored zero.
The perfect score needed a closer look at how the model won.
We cover what both results taught us about benchmark design on Oct 6.
enclave.ai/events/how-do-you…
Enclave retweeted
DeepSeek v4.1 Flash is an amazing cybersec model
nitter.cf/enclaveai/status/21001…
DeepSeek V4.1 Flash completely owned our hacking benchmark. Every vuln was discovered and exploited. And it only cost $4.
So we looked deeper:
The score looked perfect.
The traces showed why advanced agent benchmarks must check not only whether an attack worked, but how.
Full analysis: enclave.ai/blog/deepseek-v41…
Enclave retweeted
We put OpenAI's GPT-6 Sol through the AI Hacking Arena, our benchmark for verified command execution, where a working exploit counts and a written claim doesn't.
Result: 4/4 tasks, 6/11 runs verified, 0 false positives, and the fastest run on the board at 1h 06m and ~$25.
It swept Grafana with a double-URL-encoded path traversal and popped Nextcloud with PHP template injection. In Jenkins, when it couldn't escalate, it said so instead of inventing a flag.
enclave.ai/blog/gpt-6-sol-we…
We put OpenAI's GPT-6 Sol through the AI Hacking Arena, our benchmark for verified command execution, where a working exploit counts and a written claim doesn't.
Result: 4/4 tasks, 6/11 runs verified, 0 false positives, and the fastest run on the board at 1h 06m and ~$25.
It swept Grafana with a double-URL-encoded path traversal and popped Nextcloud with PHP template injection. In Jenkins, when it couldn't escalate, it said so instead of inventing a flag.
enclave.ai/blog/gpt-6-sol-we…
Enclave is now available on AWS Marketplace.
The autonomous security team for the enterprise.
Agents map your attack surface, exploit what's exploitable, and fix it, then attack again to confirm it holds.
$4.65 bought a perfect score on our hacking benchmark.
It also broke it.
Join our October 6 webinar to hear about the five attack paths we never designed and what we changed.
enclave.ai/events/how-do-you…
DeepSeek V4.1 Flash scored 11/11 on our AI hacking benchmark for $4.65, so we audited every attack path.
Six runs used the weakness the challenge was built to measure. Five found routes our test environment left open.
Scored on the six alone, the cost is 78 cents per verified run, still about 21x cheaper than the previous leader.
enclave.ai/hackingrace
Enclave will be at the CyberRisk Collaborative New York Leadership Exchange on September 30 to talk about what responsible security leadership looks like as AI systems become more autonomous.
If you’re attending, we’ll meet you there!
events.cyberriskcollaborativ…
Enclave retweeted
Well, this went viral on Hacker News...
In a new post, @Yanir_ shares more details about what we observed when testing DeepSeek V4.1 Flash on our benchmark (which we now need to extend!)
enclave.ai/blog/deepseek-v41…
Enclave retweeted
DeepSeek V4.1 Flash completely owned our hacking benchmark. Every vuln was discovered and exploited. And it only cost $4.
So we looked deeper:
The score looked perfect.
The traces showed why advanced agent benchmarks must check not only whether an attack worked, but how.
Full analysis: enclave.ai/blog/deepseek-v41…
Enclave retweeted
I agree. OpenAI can build its defense factory around Astra. But the diffusion of AI across most enterprise cybersecurity will be independent and multi-model, using the best frontier or open-weight model for each task as capabilities leapfrog.
Cognition, Factory, and Cursor already proved the pattern in coding.
Greg Brockman says OpenAI pointed Astra at its own systems until it ran out of vulnerabilities to find:
"We took 25% of our production engineers and said, 'Sorry, all your projects are on hold. You are now defending. You are now up-leveling our security architecture. You're going to use the models to find all the holes.' And we found a number of serious issues, and we fixed them."
"We found some new problems, but eventually it saturated. We basically have found, to our knowledge, all of the P0s, all of the critical problems that Astra is smart enough to find. And of course, there will be a new model, there will be a new round."
"You want to be in this tight loop of new cyber capability drops, you deploy it against your systems, you find the new holes, and ideally, you've managed to automate this, what we call defense factory. That's what we're building internally."
"There are ideas, for example, formally verifying all of software, that are possible with AI."
@gdb @bhorowitz
ICYMI: An agent swarm’s real advantage isn’t scale. It’s shared memory.
In the July a swarm tested hundreds of approaches without repeating the same work, in the Hugging Face incident.
Watch-> enclave.ai/events/defending-…
Agents will use enterprise systems at machine scale that means that security has to operate at machine scale too.
Continuously finding attack paths, repairing them, and testing them again.
Thanks for including Enclave in the conversation, @Levie.
Protecting enterprise data in a world where agents are using our systems 100X more than people ever did is going to be one of the more complex security and governance challenges of the 21st century.
Importantly, security and productivity gains are inexorably linked in the world of AI. If you give an agent too much unfettered information access, it will be difficult to truly control and protect your data; and conversely, if you lock everything down completely, you won’t get any real productivity gains from AI.
We need all new ways to protect our systems, environments, and structured and unstructured dada in the enterprise in an intelligent way by modernizing our approach to security and governance.
At Box, as one example, we’re building new intelligent ways to protect enterprise data and agentic use of that information. A recent update in Box Shield is to provide granular controls on what content agents can and can’t work with based on document classification level. We’re also working on other features that can automatically detect and alert (or block) when data is being accessed or used in unusual or anomalous ways by agents.
And this is just the start. There’s a ton more innovation coming across the entire industry - from the labs like OpenAI or Anthropic; security platforms like Palo Alto Networks, Cisco, CrowdStrike, Okta; or startups like Enclave, Method, Alterion, Runlayer, and many many others - to rethink how we protect information in the world of AI agents. Exciting and wild times ahead.
Enclave is now supported on Claude, as an MCP server.
Our autonomous security agents map every path an attacker can take through your org. They chain the small findings the way an attacker would, continuously, with no scan to schedule and no analyst to kick them off.
enclave.ai/demo