@cotoolai

Composable AI agents for security teams

Joined May 2025
We've raised a $7.4M seed round led by @a16z to build the agent operating system for security teams. Threat actors now scale with tokens. Campaigns that used to require a coordinated team can be run by a small group with the right model harness. Defense has been absorbing that hit with the same playbook and the same headcount. We built Cotool to make defense compound in the same way. Grateful to the team at @a16z , @YCombinator, @WndrCoLLC , @homebrew , and our angels from Okta, Ramp, Cloudflare, and others who've lived this problem firsthand. If you’re a security practitioner looking for more leverage in the AI age, come see how Cotool can help!
11
2
21
2,998
Today we're releasing BlueBench-Simulation, a brand new benchmark suite evaluating agent performance on defensive cybersecurity tasks against simulated enterprise environments. BlueBench-Simulation is releasing with 12 scenarios across 3 new benchmarks, each covering a different stage of the intrusion lifecycle: - Initial Access & Command-and-Control (C2) - Identity & Active Directory Attacks - Impact & Exfiltration Every scenario is a generated enterprise environment with weeks of routine activity hit by an intrusion. Models start from a single weak alert and produce a threat hunting report, an incident response report, or a detection query. - The most notable result, @SpaceXAI's Grok 4.6 led the suite at 71.8%, taking two of the three benchmarks and finishing first overall on pure capability. - @claudeai Opus 4.8 came in second at 68.9%. We'd expect Anthropic's newer models to do at least as well, but Fable 5, Opus 5, and Sonnet 5 all ran into cybersecurity refusals, so we couldn't get a clean read on them. - The @OpenAI GPT-5.6 model family once again anchored the value end of the Pareto curve More details in thread below ⬇️ Full results: cotool.ai/research
3
5
4
25
4,450
On open-weight models, Kimi K3 was the best of the group at 57.1%, fourth overall and ahead of GPT-5.6 Terra and Gemini 3.6 Flash. GLM-5.2 followed at 52.1% for $1.93 per task and holds a spot on the Pareto frontier. We're excited to follow up with results for GLM 5.3 as well!
1
2
122
BlueBench is our series of benchmarks for AI agents doing defensive security work. The Intrusion benchmarks replay real compromises. Simulation builds on that with fully generated environments: a small enterprise with normal baseline activity, an intrusion, and benign lookalikes that punish pattern-matching without verification. Because we author the environment, we're able to ensure the ground-truth is complete and that the data isn't leaked into model training sets. Our last Intrusion release, if you missed it: nitter.cf/cotoolai/status/208835…
Today we're sharing BlueBench-Intrusion-003, a new benchmark testing cyber incident response using data from a real intrusion of a live AWS environment. Each model started with a security alert and investigated events across CloudTrail, GuardDuty, S3 access logs, and VPC Flow Logs before writing a complete incident response report. Key results: 1️⃣ GPT-5.6 Sol led at 88.3% and was the most consistent top-performing model. GPT model family dominates the pareto frontier. 2️⃣ Kimi K3 led the open-weight models at 85.1%. 3️⃣ Anthropic was the only provider affected by service-level cybersecurity refusals. Full interactive results in 🧵⬇️
2
102
We @cotoolai have signed OpenAI's open letter calling for collective action on cyber defense, alongside @AnthropicAI , @Google , @CrowdStrike and other leading companies. As models get more capable, AI-enabled attacks will become far more widespread and far more sophisticated. Those same AI advances have the potential to give security and technology teams new ways to defend, fix weaknesses, and make our digital world much more secure. Cotool is committed to making AI-powered defense genuinely deployable for the teams protecting essential services, share what works instead of hoarding it, and measure ourselves on how many organizations are protected and whether the fixes hold We're privileged to stand alongside these teams in accelerating defenders' priorities with tools, funding, and hands-on support. The letter is worth reading in full: openai.com/collective-cyberd…
1
3
1
11
554
Cotool retweeted
It's been 5 weeks since we learned about the OpenAI/Hugging Face breach, which revealed an awkward reality: defenders have to ask models the same questions attackers do. In this conversation, Cotool CEO Max Pollard and Neo CEO Nick Warner sit down with a16z's Joel de la Garza at Black Hat to discuss how security tools were designed to stop people or malware (and how AI agents are neither), what breaks when half of enterprise software goes agentic, and how teams are routing around cyber refusals. 00:00 Intro 01:00 Models escaping containment & the Hugging Face breach 01:50 Why models refuse to help the good guys 05:45 Built to stop people and malware, but agents are neither 06:45 "The end justifies the means, in the mind of the model" 10:45 50% of enterprise apps agentic before 2027 14:20 When your honeypot becomes a false positive machine 15:25 Signatures are dead, and so is behavioral detection 18:55 Defending AI and defending from AI @maxpollard415 @cotoolai @neo_ai_security
20
16
5
88
35,081
Today we're sharing BlueBench-Intrusion-003, a new benchmark testing cyber incident response using data from a real intrusion of a live AWS environment. Each model started with a security alert and investigated events across CloudTrail, GuardDuty, S3 access logs, and VPC Flow Logs before writing a complete incident response report. Key results: 1️⃣ GPT-5.6 Sol led at 88.3% and was the most consistent top-performing model. GPT model family dominates the pareto frontier. 2️⃣ Kimi K3 led the open-weight models at 85.1%. 3️⃣ Anthropic was the only provider affected by service-level cybersecurity refusals. Full interactive results in 🧵⬇️
2
7
4
30
8,936
A special thanks to @ThruntingLabs for partnering on this dataset!
1
120
If your AI security operations depend on permission from another company's imperfect cyber guardrails, you have a problem. Introducing Cotool Router: a model routing layer built for security agents. Router delivers frontier-quality results with 0 cyber-refusals for defensive security tasks. Router will detect when a chosen frontier model hits cyber refusals and automatically route to the best substitute, meaning zero interruption for your agents.
1
4
15
1,525
Performance is not the only gap. The wrong model choice can cost nearly four times more and perform worse on the same security workload. Cotool Router provides evaluation-based model selection so teams can pick the optimal model for every task.
1
3
85
Customers using Router today are seeing 100% task completion across routed defensive security workloads. Defenders need to scale AI in their systems now more than ever. Cotool Router helps security teams operate with resilience and is available now. Read the full announcement: cotool.ai/blog/introducing-c…
4
66
GPT 5.6 dropped this morning and it now completely dominates the pareto frontier of our newest security operations benchmark task, BlueBench-Intrusion-002. BlueBench-Intrusion-002 is a real Windows enterprise intrusion, 4M log events, 22 models tested across detection engineering, malware analysis, IR, and threat hunting. This is the second release in BlueBench, @cotoolai's benchmark suite focused purely on defensive security operations tasks. Most public security evals cover code vulnerability patching or CTF puzzles in toy environments. BlueBench tasks are built with real intrusion data in realistic eval environments. Details on the results in thread ⬇️
1
5
2
24
13,164
The other headline: this is the first time we've seen open weights reach the frontier on realistic defensive security work. @Zai_org GLM-5.2 sits at 72.5%, on par with GPT-5.5 and just shy of Opus 4.7/4.8. @Kimi_Moonshot Kimi K2.6 holds a spot on the cost frontier, as well. That said, the "frontier" moves quickly and with the release of GPT 5.6, we see the gap between closed lab and open weight models widen back again. Check out full details and interactive results here: cotool.ai/research/windows-e…
1
1
5
512
Huge shoutout to @ThruntingLabs for being a fantastic partner is this eval!
2
4
870
Go team USA 🇺🇸🇺🇸🇺🇸
2
156