@dreadnode

Where security agents run. AI infrastructure to build, evaluate, and deploy with confidence.

Joined August 2010
Worried about your production agents going out of scope? Us too. AgentJudge is our agent hall monitor that stops out of scope tool calls before they execute. Before a tool runs, the judge reads the agent’s intent, the proposed call, and a rubric you define, then returns allow, deny, or ask (escalate to you). The agent stays autonomous; the judge is the guardrail. Available in the TUI today, UI updates coming to the Dreadnode Platform soon! 👀 Get Started: docs.dreadnode.io/getting-st… AgentJudge Docs: docs.dreadnode.io/tui/guard-… Related Research: dreadnode.io/research/scope-…
9
1
26
3,314
dreadnode retweeted
scopejudge 🤝 jev
Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]
1
1
5
618
Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]
2
10
3
36
3,329
Jev adds a fast check for contextual scope decisions, a promising step toward efficient runtime judges. Of course, not everything is a nail with this new hammer. Hard limits belong in code: permissions, network restrictions, and sandbox controls. If code can decide, enforce it there. [4/4]
1
1
198
👀👀👀👀👀👀👀👀👀👀
Replying to @typesafeai
@typesafeai 's Jev definitely earns its hype. Results soon from experiments we've been up to @dreadnode
1
1
1
10
1,163
Presenting at @labscon_io is always a homecoming. The conference that gave me my start. And I am so happy I got to speak at the last one. We had a lot of fun Mogging Mythos and showing how to push the offensive capabilities of open-weight models to their edge and beyond.
1
4
14
859
We're out here cybermaxxing models and mogging Mythos. Thanks to all who attended @Dr_Machinavelli's @LabsSentinel LabsCon talk this afternoon 😎
Martin Wendiggensen (@Dr_Machinavelli) closing out the morning keynotes with: Why Flexing Offensive Muscles Teaches Us How To Defend In The Age Of AI
1
2
7
557
💪💪💪💪💪 Tomorrow (9/17) @Dr_Machinavelli takes the stage at the final @SentinelOne @labscon_io to discuss how to leverage offensive cyber capabilities to improve defenses in the age of AI. More info: labscon.io/speakers/martin-w…
2
241
Qwen 3.8 Flash eval results are live on DreadIndex, landing at #19 on our leaderboard. It does well for its cost, but remains light on offensive security capability (not surprising given it is a flash model). See how it compares to other models: dreadnode.io/research/dreadi…
1
4
280
METR is just one of many existing AI evaluators! If you take issue with METR, that isn’t a good argument against requiring companies to carry out embedded audits Here are 20 other evaluators who work with AI labs to assess risk: Transluce Grey Swan Apollo AVERI SecureBio Faculty Vaultis Dreadnode Irregular RAND Far AI ActiveFence (now Alice) Active Site Deloitte (via Gryphon acquisition) Nemesys Mercor (via Sepal acquisition) AE Studio Scale Frontier Design Redwood
13
28
8
280
16,835
GLM-5.3-Flash results now on DreadIndex: dreadnode.io/research/dreadi…
6
17
1,189
Here's where to catch the Dreadnode crew at year two of @OffensiveAIcon: > Join us for the welcome reception at The Shelter Club on Sunday evening! > @mkultraWasHere is closing out Day One of talks, presenting on model cheating behavior. > Dynamic duo @shanejcaldwell and @0xdab0 take the stage on Tuesday for a session on implementing a judge model as a runtime monitor, and how to keep agents in scope. See you in Oceanside! 🏄
1
6
1
11
634
Always feels a little strange to watch a recording of myself but had a lot of fun repping @dreadnode and digging into the AI cyber nexus with the @AtlanticCouncil’s @CyberStatecraft crew.
“The US hasn’t demonstrated a willingness to take the lead” on data security and privacy, says @Dreadnode Head of Policy @velvethamm3r. With AI accelerating progress, “we’re now seeing real-time consequences,” she tells @CyberStatecraft’s Trey Herr.
1
4
20
2,358
dreadnode retweeted
HYBRID 8 SEP 13:00 UTC - Data Governance for AI and Healthcare: From Washington to Beijing isoc.live/21983 @CyberStatecraft convenes @Pfizer's Ranjit Kumble, @Snowflake's Stephen Moon, @dreadnode's Daria Bahrami and Harvard Medical School's Gil Alterovitz on the data, compute and regulation behind AI-enabled healthcare. Moderated by @GrahamBrookie and Trey Herr. In DC and online. #ACCyber #AI #HealthData #DataGovernance
1
1
137
dreadnode retweeted
It's not every day your research goes to prod, but it's always a good day when it does. Thanks to @mkultraWasHere for making it happen. More monitoring features to come as @dreadnode swarms on making sure your swarms only swarm where you want them to. Back to work 🫡
Worried about your production agents going out of scope? Us too. AgentJudge is our agent hall monitor that stops out of scope tool calls before they execute. Before a tool runs, the judge reads the agent’s intent, the proposed call, and a rubric you define, then returns allow, deny, or ask (escalate to you). The agent stays autonomous; the judge is the guardrail. Available in the TUI today, UI updates coming to the Dreadnode Platform soon! 👀 Get Started: docs.dreadnode.io/getting-st… AgentJudge Docs: docs.dreadnode.io/tui/guard-… Related Research: dreadnode.io/research/scope-…
2
1
7
598