@WolfBenchAIi
iAccount based inGermany
About this account
- Account based in
- Germany
- Connected via
- Germany Android App
Account-level information from X, not a live location or the device used for a specific post.
https://nitter.cf/t.co/IGlGiHLlnE // @WolframRvnwlf's new evaluation framework for models and agents: because one score is not enough! // brought to you by @CoreWeave/@wandb
Joined March 2026
- Tweets37
- Following15
- Followers334
- Likes142
Pinned Tweet
WolfBench results for @Zai_org's new GLM-5.2:
#1 open-weight model
#3 overall
Cheapest among the top 10
A massive leap forward for open models - and for capable, affordable intelligence you can truly own. ๐
GPT-5.6 Sol took the top three spots on WolfBench:
Codex: 86.74%
Terminus-2: 85.17%
Hermes: 84.49%
So far, so benchmarky. Then I checked the tokens.
With the same Sol model at maximum reasoning, Codex used 83.8 million tokens per run. Hermes used 170.9 million. Twice the tokens for a 2.25 percentage-point lower score.
Hermes recorded $173.41 per run. Codex comes out at an estimated ~$89.79-$95.06. (That Codex number is reconstructed from token usage, not billed spend, but even its upper bound is 45% lower.)
The GPT-5.5 comparison makes it even weirder: With Codex, GPT-5.6 gained 8.09 points, used 20% fewer tokens, and appears roughly half as expensive. With Terminus-2, it gained 8.31 points while cost fell 26.5%. With Hermes, it gained 10.34 points - but used 68% more tokens and cost 47% more.
So is GPT-5.6 cheaper than GPT-5.5? Depends entirely on the agent.
One last oddity: Across 15 Sol runs, the "configure git webserver" task was never solved. Terra solved it once. Luna solved it every time. The smaller variants are not simply weaker copies of Sol.
My takeaway: Codex + Sol is the strongest full-agent stack we tested. Terminus-2 + Sol is the cleanest reference baseline. And for Hermes, I would seriously consider Terra: only 3.6 points behind Sol, but 44% cheaper.
Benchmark the stack, not just the model! And run it more than once.
Full results on WolfBench.ai:
Replying to @WolfBenchAI @Zai_org
WolfBench.ai is not just a leaderboard - it is an interactive results viewer.
Enable 3D Tokens + Cost view (tokens = depth, cost = shadow) to see efficiency at a glance: top 3 performance with low token use and lowest spend.
It is the cheapest model in the top 10+.
Most benchmark charts show only average scores, but WolfBench's Solid Base metric shows how consistent a model truly is: the percentage of tasks it solved in EVERY one of the 5 runs. By that measure, GLM-5.2 ranks #1, solving more tasks consistently than any other model I tested.
We spent $11,081.12 evaluating @AnthropicAI's Claude Fable 5 on WolfBench.
Our most expensive benchmark yet.
And it did not even top the charts.
Not because it lacked capability, but because it kept refusing.
Details in thread: ๐งต
Even without refusals, real agentic failure patterns remained. Fable's biggest weakness was overconfident self-verification: it often declared victory too early once the solution looked plausible, while the actual benchmark checks still caught wrong output, messy cleanup, missed edge cases, or slow code.
Claude Fable 5 may be exceptional, but it is currently not the best fit as a general-purpose agentic daily driver. It is too expensive and too refusal-prone to turn its strengths into efficient, reliable agentic work.
Explore the full results at WolfBench.ai: compare models and agents in the interactive chart, and click any bar to open the corresponding evals and traces in @weave_wb for deeper inspection.
For benchmarks, I keep agent versions stable so results stay comparable. But new models can expose agent-side bugs. Here, updating @openclaw from 2026.3.11 to 2026.4.23 lifted Kimi K2.6 from 4% to 60% on @WolfBenchAI due to crucial fixes in how the agent handles its tool calling.
GPT-5.5 takes over WolfBench! Itโs now the #1 model, ahead of Claude Opus 4.7 and 4.6, GPT-5.4, Sonnet 4.6, Kimi K2.6, Gemini 3.1 Pro, and more.
Notable findings after 30 runs (40h runtime, >1.7B tokens, ~$3K cost):
- @OpenAI's GPT-5.5 is the best model we ever tested.
- @cursor_ai's Agent CLI (CA) is the best agent we ever tested.
- @NousResearch's Hermes Agent (HA) outperformed OpenClaw (OC).
- With Hermes, going from medium to xhigh reasoning only improved consistency, not capability.
Note: This is WolfBench, where we look at more than just the average score, because one metric is not enough. The golden โ
score is the actual 5-run average, which most other benchmarks report as their only score. โ
shows the ceiling (what percentage of the full benchmark this model+agent combination solved at least once across all runs). โ shows the solid base (what percentage of the full benchmark it solved consistently in every run).
Let's compare the WolfBench top model, GPT-5.5, with our #2, Claude Opus 4.7:
- @openclaw is still better on Opus 4.7 than on GPT-5.5: 75% vs. 70%, with a slightly higher ceiling and base. This is also the third-highest score across all models and agents - only @cursor_ai and Terminus-2 (the official @terminalbench 2.0 test harness) rank higher, both at 77%.
- When no reasoning level is set, the OpenAI API defaults GPT-5.5 to medium, while the Anthropic API defaults Opus 4.7 to no thinking. That's why Terminus-2 and Hermes Agent have different effort levels.
- Note that higher effort levels don't necessarily improve scores in agentic benchmarks - thinking harder can actually make the model dumber: wandb.ai/wandb_fc/wolfbench-โฆ
- Still have to evaluate Cursor with Opus 4.7; with 4.6, it got 63%.
Replying to @WolfBenchAI @OpenAI
Visit wolfbench.ai for the full lineup of models and agents. The site is fully interactive: filter and sort by models, agents, metrics, and scores, or click any bar to jump straight to the corresponding @weave_wb evals and traces. Full transparency for all 300+ runs!
WolfBench retweeted
the super interesting thing that I find not enough people talking about is OpenClaw topping the T2 leaderboard for Opus 4.7 with thinking off (@WolfBenchAI eval harness) l - OC also a generic harness unlike the other harnesses in the below lb which are likely benchmaxxed for coding
WolfBench retweeted
Replying to @Aiolias_ @WolfBenchAI
That's fair.
But this one is a bit different and tells a realistic story (my custom testing pipeline share more than half of what it uses) .
WolfBench retweeted
๋ฉฐ์น ์ ๋ถํฐ ์๊พธ Hermes ์์ด์ ํธ์ ์ ๊ฒฝ์ด ์ฐ์ธ๋ค.
์ฌ์ค OpenClaw๊ฐ ์ข ๋ ์ค๋ ์์ฅ์ ์ฅ์
ํ ์ค ์์๋๋ฐ, ์์ง ๊ฒ์ฆ์ ์๋์ง๋ง ๊ฐ๋ ฅํ ๊ฒฝ์์๊ฐ ๋ค์ด์จ ๊ฒ ๊ฐ๋ค.
๋ฏธ๊ตญ์ NousResearch๋ผ๋ ํ์ด ์๋ค. Nous Research๋ ์คํ์์ค AI ๋ถ์ผ์์ ๊ฐ์ฅ ์์๊ฐ๋ ์คํํธ์
/์ฐ๊ตฌ ํ ์ค ํ๋์ด๊ณ . ์ฌ์ฉ์๊ฐ ์ง์ ์ ์ดํ ์ ์๋ โuser-aligned(์ฌ์ฉ์ ์ ๋ ฌ)โ ๋ชจ๋ธ๋ก ํฐ ์ฃผ๋ชฉ์ ๋ฐ๊ณ ์๋ค.
๊ทธ๋ค์ด ๋ง๋ Hermes Agent๊ฐ 89๊ฐ ์ค์ ์์
ํ
์คํธ์์ Claude Code์ OpenClaw๋ฅผ ์์ง๋ ๋ค. ์ ์๋ง ๋์ ๊ฒ ์๋๋ผ "๋ฐ๋ฅ"์ด ๋์๋ค. ๋งค๋ฒ ๋ ๋ง์ ์์
์ ์์ ์ ์ผ๋ก ์๋ฃํ๋ค๋ ๋ป์ด๋ค.
๊ทธ๋ผ ์ ๊ทธ๋ฐ ๊ฒฐ๊ณผ๊ฐ ๋์์๊น? ํต์ฌ์ ํ๋ค์ค๋ค.
ํ๋ค์ค๋ AI ๋ชจ๋ธ์ ๊ฐ์ธ๋ ํ์ด๋ค. ๊ฐ์ Opus 4.6์ด๋ผ๋ ์ด๋ค ํ๋ค์ค์ ๋ฃ๋๋์ ๋ฐ๋ผ ๊ฒฐ๊ณผ๊ฐ ๋ฌ๋ผ์ง๋ค. Hermes์ ์ฃผ์ฅ์ "์ฐ๋ฆฌ ๋ชจ๋ธ์ด ๋ ์ข๋ค"๊ฐ ์๋๋ค. "๊ฐ์ ๋ชจ๋ธ์ ๋ ์ ์ฐ๋ ๊ตฌ์กฐ๋ฅผ ๋ง๋ค์๋ค"๋ ๊ฑฐ์ ์๋ฏธ๊ฐ ์๋ค.
๊ทธ ๊ตฌ์กฐ์ ํต์ฌ์ ํ์ต ๋ฃจํ๋ผ๋ ํต์ฌ ๊ธฐ์ ์ด๋ค.
Claude Code๋ ๋งค๋ฒ ์๋ก ์์ํ๋ค. OpenClaw๋ MEMORY.md๋ก ๊ธฐ์ต์ ์๋ ๊ด๋ฆฌํ๋ค. ๊ธฐ์ต์ ์ ์งํ๊ฒ ์
ํ
์ ํ๋ ๊ฒ์ ์ฌ์ ํ ์ธ๊ฐ ๋ชซ์ด๋ค.
Hermes๋ ์์คํ
์ด ์กฐ๊ธ ๋ค๋ฅด๋ค. ๋ณต์กํ ์์
์ด ๋๋๋ฉด ์์ด์ ํธ๊ฐ ์์จ์ ์ผ๋ก ์ฌ์ฌ์ฉ ๊ฐ๋ฅํ ์คํฌ์ ์์ฑํ๊ณ ์ ์ฅํ๋ค. ๋ญ ๊ธฐ์ตํ ์ง, ๋ญ ์คํฌ๋ก ๋ง๋ค์ง ์์ด์ ํธ๊ฐ ์ค์ค๋ก ํ๋จํ๋ ๊ตฌ์กฐ๋ค.
๊ณต์ ์ค๋ช
๊ทธ๋๋ก "built-in learning loop", "autonomous skill creation", "skills self-improve during use." ์ธ๊ฐ์ด ์๋ฌด๊ฒ๋ ์ ํด๋ ์์ด์ ํธ๊ฐ ์ ์ ์๋ฆฌํด์ง๋ค.
์ปค๋ฎค๋ํฐ์์๋ Hermes๋ฅผ "Claude Code ์คํ์ผ CLI์ OpenClaw ์คํ์ผ ๋ฉ์์ง ์์ด์ ํธ์ ์ค๊ฐ"์ผ๋ก ๋ถ๋ฅด๊ธฐ๋ ํ๋ค. ๋ ๋ค ๋๋ ค ํ๋ค๋ ๋ป์ด๋ค. ํฐ๋ฏธ๋์์๋, ํ
๋ ๊ทธ๋จ์์๋, VPS์์๋. v0.2.0 ์ถ์ ์ดํ ๋น ๋ฅด๊ฒ ์คํ 10,000๊ฐ๋ฅผ ๋๊ฒผ๊ณ , ํ์ฌ 22,000๊ฐ๋ฅผ ๋ํํ๋ค.
๊ฐ์ธ์ ์ผ๋ก๋ ์ด ์ ๋ต์ ์๋ฆฌํ๋ค๊ณ ์๊ฐํ๋ค. ์ฌ๋๋ค์ "3% ๋ ๋๋ํ ๋ชจ๋ธ"๋ณด๋ค "๋๋ฅผ ๊ธฐ์ตํ๋ ์์ด์ ํธ"๋ผ๋ ์คํ ๋ฆฌ์ ๋ ๋๋ฆฐ๋ค. Hermes์ ์ฌ๋ก๊ฑด "The agent that grows with you"๋ ์ฑ๋ฅ์ด ์๋๋ผ ๋์ ์์ด์ ํธ์ ๊ด๊ณ๋ฅผ ํ๋ค.
AI ์์ด์ ํธ์ ๋ค์ ์ ์ํฐ๋ ๋ชจ๋ธ ์ฑ๋ฅ์ด ์๋๋ค. ์ผ๋ง๋ ๋น ๋ฅด๊ฒ ๋ฐฐ์ฐ๊ณ , ์ผ๋ง๋ ์ค๋ ๊ธฐ์ตํ๋๋๊ฐ ์๋๊น?
๊ทธ๋ฆฌ๊ณ ๊ฐ์ธ ๋ง์ถค ์์ด์ ํธ ๋ธ๋๋๊ฐ ์ ์ ๋ค๊ฐ์ค๋ ๋๋์ด๋ค.
๋๋ ์ค๋ ํ๋ฒ ์ค์นํ๊ณ ๋๋ ค๋ณด๋ ค๊ณ ํ๋ค.
Hermes Agent outperformed Claude Code and OpenClaw as an agentic harness for both Opus 4.6 and GPT-5.4 on 89 real-world tasks.
Not just higher scores but a higher floor. More tasks solved reliably, every single run.
@Teknium1 @NousResearch really cooked with this one. ๐ฅ
We just published our evals for Agent Harnesses on WolfBench and Hermes out of the box came out on top.
nitter.cf/WolfBenchAI/status/203โฆ
Hermes Agent outperformed Claude Code and OpenClaw as an agentic harness for both Opus 4.6 and GPT-5.4 on 89 real-world tasks.
Not just higher scores but a higher floor. More tasks solved reliably, every single run.
@Teknium1 @NousResearch really cooked with this one. ๐ฅ
Hermes Agent outperformed Claude Code and OpenClaw as an agentic harness for both Opus 4.6 and GPT-5.4 on 89 real-world tasks.
Not just higher scores but a higher floor. More tasks solved reliably, every single run.
@Teknium1 @NousResearch really cooked with this one. ๐ฅ
Key takeaways from our latest eval:
> Hermes Agent (default settings) hits 64% avg on Opus 4.6 vs Claude Code's 63% and OpenClaw's 58% โ but the solid base jumps from 45%/42% to 49%.
> On GPT-5.4 the gap is massive: 66% avg vs Claude Code's 48% and OpenClaw's 61%, solid base 47% vs 22%/45%.
> It takes GPT-5.4 with xhigh effort for OpenClaw to surpass Hermes Agent with default=medium effort.
> Only with xhigh effort could OC surpass HA with GPT-5.4.
Full breakdown here: wolfbench.ai