@dreadnodei
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Where security agents run. AI infrastructure to build, evaluate, and deploy with confidence.
Joined August 2010
- Tweets421
- Following123
- Followers3K
- Likes274
Pinned Tweet
Worried about your production agents going out of scope? Us too. AgentJudge is our agent hall monitor that stops out of scope tool calls before they execute.
Before a tool runs, the judge reads the agent’s intent, the proposed call, and a rubric you define, then returns allow, deny, or ask (escalate to you).
The agent stays autonomous; the judge is the guardrail.
Available in the TUI today, UI updates coming to the Dreadnode Platform soon! 👀
Get Started: docs.dreadnode.io/getting-st…
AgentJudge Docs: docs.dreadnode.io/tui/guard-…
Related Research: dreadnode.io/research/scope-…
dreadnode retweeted
scopejudge 🤝 jev
Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does.
@typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]
Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does.
@typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]
Jev adds a fast check for contextual scope decisions, a promising step toward efficient runtime judges.
Of course, not everything is a nail with this new hammer. Hard limits belong in code: permissions, network restrictions, and sandbox controls. If code can decide, enforce it there. [4/4]
About ScopeJudge: dreadnode.io/research/scope-…
ScopeJudge on GitHub: github.com/dreadnode/scopeju…
👀👀👀👀👀👀👀👀👀👀
Replying to @typesafeai
@typesafeai 's Jev definitely earns its hype. Results soon from experiments we've been up to @dreadnode
dreadnode retweeted
Presenting at @labscon_io is always a homecoming. The conference that gave me my start. And I am so happy I got to speak at the last one. We had a lot of fun Mogging Mythos and showing how to push the offensive capabilities of open-weight models to their edge and beyond.
We're out here cybermaxxing models and mogging Mythos. Thanks to all who attended @Dr_Machinavelli's @LabsSentinel LabsCon talk this afternoon 😎
Martin Wendiggensen (@Dr_Machinavelli) closing out the morning keynotes with:
Why Flexing Offensive Muscles Teaches Us How To Defend In The Age Of AI
💪💪💪💪💪 Tomorrow (9/17) @Dr_Machinavelli takes the stage at the final @SentinelOne @labscon_io to discuss how to leverage offensive cyber capabilities to improve defenses in the age of AI.
More info: labscon.io/speakers/martin-w…
Qwen 3.8 Flash eval results are live on DreadIndex, landing at #19 on our leaderboard.
It does well for its cost, but remains light on offensive security capability (not surprising given it is a flash model). See how it compares to other models: dreadnode.io/research/dreadi…
dreadnode retweeted
METR is just one of many existing AI evaluators! If you take issue with METR, that isn’t a good argument against requiring companies to carry out embedded audits
Here are 20 other evaluators who work with AI labs to assess risk:
Transluce
Grey Swan
Apollo
AVERI
SecureBio
Faculty
Vaultis
Dreadnode
Irregular
RAND
Far AI
ActiveFence (now Alice)
Active Site
Deloitte (via Gryphon acquisition)
Nemesys
Mercor (via Sepal acquisition)
AE Studio
Scale
Frontier Design
Redwood
GLM-5.3-Flash results now on DreadIndex: dreadnode.io/research/dreadi…
Here's where to catch the Dreadnode crew at year two of @OffensiveAIcon:
> Join us for the welcome reception at The Shelter Club on Sunday evening!
> @mkultraWasHere is closing out Day One of talks, presenting on model cheating behavior.
> Dynamic duo @shanejcaldwell and @0xdab0 take the stage on Tuesday for a session on implementing a judge model as a runtime monitor, and how to keep agents in scope.
See you in Oceanside! 🏄
More details: dreadnode.io/company/newsroo…
dreadnode retweeted
Always feels a little strange to watch a recording of myself but had a lot of fun repping @dreadnode and digging into the AI cyber nexus with the @AtlanticCouncil’s @CyberStatecraft crew.
“The US hasn’t demonstrated a willingness to take the lead” on data security and privacy, says @Dreadnode Head of Policy @velvethamm3r.
With AI accelerating progress, “we’re now seeing real-time consequences,” she tells @CyberStatecraft’s Trey Herr.
dreadnode retweeted
HYBRID 8 SEP 13:00 UTC - Data Governance for AI and Healthcare: From Washington to Beijing
isoc.live/21983
@CyberStatecraft convenes @Pfizer's Ranjit Kumble, @Snowflake's Stephen Moon, @dreadnode's Daria Bahrami and Harvard Medical School's Gil Alterovitz on the data, compute and regulation behind AI-enabled healthcare. Moderated by @GrahamBrookie and Trey Herr. In DC and online.
#ACCyber #AI #HealthData #DataGovernance
ALT Dark blue cybersecurity-themed title card for the Atlantic Council Cyber Statecraft Initiative. A stethoscope rests over a glowing circuit-board pattern, visually linking healthcare and digital technology. Large yellow and white text reads “Data Governance for AI and Healthcare: From Washington to Beijing,” with the date 8 September 2026 below. Atlantic Council branding appears at upper left, with #ACCYber and partner logos along the bottom.
dreadnode retweeted
It's not every day your research goes to prod, but it's always a good day when it does.
Thanks to @mkultraWasHere for making it happen. More monitoring features to come as @dreadnode swarms on making sure your swarms only swarm where you want them to.
Back to work 🫡
Worried about your production agents going out of scope? Us too. AgentJudge is our agent hall monitor that stops out of scope tool calls before they execute.
Before a tool runs, the judge reads the agent’s intent, the proposed call, and a rubric you define, then returns allow, deny, or ask (escalate to you).
The agent stays autonomous; the judge is the guardrail.
Available in the TUI today, UI updates coming to the Dreadnode Platform soon! 👀
Get Started: docs.dreadnode.io/getting-st…
AgentJudge Docs: docs.dreadnode.io/tui/guard-…
Related Research: dreadnode.io/research/scope-…