Where security agents run. AI infrastructure to build, evaluate, and deploy with confidence.

Joined August 2010
Include

Only show posts containing:

Exclude

Hide posts containing:

Time range
-
Minimum likes
dreadnode retweeted
scopejudge 🤝 jev
dreadnode
@dreadnode
Sep 21
Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]
1
1
5
618
dreadnode
@dreadnode
Sep 21
2
186
dreadnode
@dreadnode
Sep 21
Jev adds a fast check for contextual scope decisions, a promising step toward efficient runtime judges. Of course, not everything is a nail with this new hammer. Hard limits belong in code: permissions, network restrictions, and sandbox controls. If code can decide, enforce it there. [4/4]
1
1
198
dreadnode
@dreadnode
Sep 21
Then, we tested judging escalation: start with Jev and bring in smarter judges when needed. Low-confidence decisions go to GLM; disagreements go to Opus. Most decisions stayed with Jev. The chain scored comparably to the paper’s best at half the cost. [3/4]
1
1
236
dreadnode
@dreadnode
Sep 21
We tested Jev on 4,897 recorded agent actions, compared its decisions with human reviewers, and measured it against the paper’s published results. When provided the user’s request and an action to check, Jev had the highest score in that setting. [2/4]
1
1
271
dreadnode
@dreadnode
Sep 21
Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]
2
10
3
36
3,329
dreadnode
@dreadnode
Sep 18
👀👀👀👀👀👀👀👀👀👀
Replying to @typesafeai
@typesafeai 's Jev definitely earns its hype. Results soon from experiments we've been up to @dreadnode
1
1
1
10
1,163
dreadnode
@dreadnode
Sep 17
We're out here cybermaxxing models and mogging Mythos. Thanks to all who attended @Dr_Machinavelli's @LabsSentinel LabsCon talk this afternoon 😎
Martin Wendiggensen (@Dr_Machinavelli) closing out the morning keynotes with: Why Flexing Offensive Muscles Teaches Us How To Defend In The Age Of AI
1
2
7
557
dreadnode
@dreadnode
Sep 16
💪💪💪💪💪 Tomorrow (9/17) @Dr_Machinavelli takes the stage at the final @SentinelOne @labscon_io to discuss how to leverage offensive cyber capabilities to improve defenses in the age of AI. More info: labscon.io/speakers/martin-w…
2
241
dreadnode
@dreadnode
Sep 16
Qwen 3.8 Flash eval results are live on DreadIndex, landing at #19 on our leaderboard. It does well for its cost, but remains light on offensive security capability (not surprising given it is a flash model). See how it compares to other models: dreadnode.io/research/dreadi…
1
4
280
dreadnode
@dreadnode
Sep 10
GLM-5.3-Flash results now on DreadIndex: dreadnode.io/research/dreadi…
6
17
1,189
dreadnode
@dreadnode
Sep 9
Here's where to catch the Dreadnode crew at year two of @OffensiveAIcon: > Join us for the welcome reception at The Shelter Club on Sunday evening! > @mkultraWasHere is closing out Day One of talks, presenting on model cheating behavior. > Dynamic duo @shanejcaldwell and @0xdab0 take the stage on Tuesday for a session on implementing a judge model as a runtime monitor, and how to keep agents in scope. See you in Oceanside! 🏄
1
6
1
11
634
dreadnode
@dreadnode
Sep 3
Worried about your production agents going out of scope? Us too. AgentJudge is our agent hall monitor that stops out of scope tool calls before they execute. Before a tool runs, the judge reads the agent’s intent, the proposed call, and a rubric you define, then returns allow, deny, or ask (escalate to you). The agent stays autonomous; the judge is the guardrail. Available in the TUI today, UI updates coming to the Dreadnode Platform soon! 👀 Get Started: docs.dreadnode.io/getting-st… AgentJudge Docs: docs.dreadnode.io/tui/guard-… Related Research: dreadnode.io/research/scope-…
9
1
26
3,314
dreadnode
@dreadnode
Sep 2
Dreadnode side quest: ALFRED (Agentic Latex for Research, Editing, and Drafting) Principal AI Research Engineer @mkultraWasHere built a helpful LaTeX agent to support research writing — and today we’re open-sourcing it. You describe the paper, it sets up the template, pulls citations, and builds the PDF framework. Conference templates, lit and peer reviews, bring your own model, everything runs locally. Watch this tutorial for a tour of the agent, from install through first compiled draft: youtube.com/watch?v=ZP0Nnyvo… Repo: github.com/dreadnode/alfred
1
7
1
26
3,322
dreadnode
@dreadnode
Sep 1
Read more on the US open-weight model performance gap via our Head of Policy @velvethamm3r: dreadnode.io/research/from-c…
2
267
dreadnode
@dreadnode
Sep 1
Two new additions to #DreadIndex: nemotron-3-ultra-550b-a55b and gemma-4-31b-it. Landing at the bottom of the leaderboard, these evaluations offer two more proof points to increase investment in US open models. 🔗: dreadnode.io/research/dreadi… NEW: we recently added a toggle to view only open weight models.
5
7
1
14
2,076
dreadnode
@dreadnode
Aug 27
New article from @CNET discusses why AI agents keep hacking their way out of test environments — and what to do about it. Dreadnode Principal Research Engineer @shanejcaldwell's answer: an AI hall monitor. Agent judges for runtime monitoring, escalating to a human when scope is breached. Read the full article: cnet.com/tech/services-and-s…
1
5
1
12
2,016
dreadnode
@dreadnode
Aug 27
Mine The Gap Heading to Vegas for @CrowdStrike Fal.Con next week? Don't miss @Dr_Machinavelli's Day Zero keynote where he pits red and blue team agents against each other to generate training data that can be used to increase AI performance, impose realistic constraints, and operate at scale. 🗣️ CrowdStrike Day Zero Threat Research Summit 📍 Virgin Hotel Las Vegas 🗓️ Monday, August 31 ⌚ 9:15 - 9:45 AM PT 🔗 crowdstrike.com/en-us/events…
2
3
8
596
dreadnode
@dreadnode
Aug 26
New models are regularly being added to DreadIndex: dreadnode.io/research/dreadi… Observations from the latest evals—Deepseek v4 Pro 0813, GLM 5.3, and Qwen 3.8 Max—in the thread 🧵⬇️ Have you tried these models out yet? Curious if our eval results align to first-hand operator usage.
1
5
1
14
1,097