Where security agents run. AI infrastructure to build, evaluate, and deploy with confidence.

Joined August 2010
Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]
2
10
3
36
3,333
💪💪💪💪💪 Tomorrow (9/17) @Dr_Machinavelli takes the stage at the final @SentinelOne @labscon_io to discuss how to leverage offensive cyber capabilities to improve defenses in the age of AI. More info: labscon.io/speakers/martin-w…
2
241
Qwen 3.8 Flash eval results are live on DreadIndex, landing at #19 on our leaderboard. It does well for its cost, but remains light on offensive security capability (not surprising given it is a flash model). See how it compares to other models: dreadnode.io/research/dreadi…
1
4
280
GLM-5.3-Flash results now on DreadIndex: dreadnode.io/research/dreadi…
6
17
1,189
Here's where to catch the Dreadnode crew at year two of @OffensiveAIcon: > Join us for the welcome reception at The Shelter Club on Sunday evening! > @mkultraWasHere is closing out Day One of talks, presenting on model cheating behavior. > Dynamic duo @shanejcaldwell and @0xdab0 take the stage on Tuesday for a session on implementing a judge model as a runtime monitor, and how to keep agents in scope. See you in Oceanside! 🏄
1
6
1
11
634
Worried about your production agents going out of scope? Us too. AgentJudge is our agent hall monitor that stops out of scope tool calls before they execute. Before a tool runs, the judge reads the agent’s intent, the proposed call, and a rubric you define, then returns allow, deny, or ask (escalate to you). The agent stays autonomous; the judge is the guardrail. Available in the TUI today, UI updates coming to the Dreadnode Platform soon! 👀 Get Started: docs.dreadnode.io/getting-st… AgentJudge Docs: docs.dreadnode.io/tui/guard-… Related Research: dreadnode.io/research/scope-…
9
1
26
3,316
Two new additions to #DreadIndex: nemotron-3-ultra-550b-a55b and gemma-4-31b-it. Landing at the bottom of the leaderboard, these evaluations offer two more proof points to increase investment in US open models. 🔗: dreadnode.io/research/dreadi… NEW: we recently added a toggle to view only open weight models.
5
7
1
14
2,076
Mine The Gap Heading to Vegas for @CrowdStrike Fal.Con next week? Don't miss @Dr_Machinavelli's Day Zero keynote where he pits red and blue team agents against each other to generate training data that can be used to increase AI performance, impose realistic constraints, and operate at scale. 🗣️ CrowdStrike Day Zero Threat Research Summit 📍 Virgin Hotel Las Vegas 🗓️ Monday, August 31 ⌚ 9:15 - 9:45 AM PT 🔗 crowdstrike.com/en-us/events…
2
3
8
596
Definite improvement from Qwen 3.8 Max over 3.7 Max, but the Qwen series still hasn't cracked top 10 on DreadIndex due to its poor network ops performance.
136