@bridgebenchi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
The official AI benchmark of the vibe coding movement @bridgemindai
United States
Joined March 2026
- Tweets1.6K
- Following11
- Followers12.1K
- Likes1.6K
Bridgebench retweeted
Claude Sonnet 5.5 created this video entirely with code.
Sonnet 5.5 Max effort is better than Claude Opus 5.5.
Bridgebench retweeted
Claude Sonnet 5.5 just ONE SHOT Mario Kart.
This is by far the best result that I have seen on this test.
Bridgebench retweeted
Claude Sonnet 5.5 just ONE SHOT this zombies game.
Max effort is noticeably better than every other level but it is expensive.
This game cost $177 and took 49 minutes to build.
This is insane.
Claude Sonnet 5.5 just dropped.
We put it through the BridgeBench black hole merger test against Opus 5.5 and Fable 5.1.
Claude Fable 5.1: $1.22, 4m 44s
Claude Opus 5.5: $0.69, 6m 8s
Claude Sonnet 5.5: $0.34, 4m 14s
Sonnet 5.5 was the fastest and cheapest of the three.
Half the cost of Opus 5.5.
Which one did the best?
Bridgebench retweeted
Claude Sonnet 5.5 is imminent.
Sonnet 5 was a disaster. Terrible tokenizer, wasted tokens, F tier in my rankings. I never used it.
But I am hearing that Sonnet 5.5 will be the same leap that Opus 5 had with Opus 5.5.
Sonnet 5.5 is going to be a leap.
A lot of people think Claude and Codex are silently cutting their usage limits.
Right now there's no way to prove it.
So we're building UsageBench on BridgeBench.
We track the 5 hour and weekly limits over time.
If a lab ever cuts your limits in the background, you'll see it here first.
Have you noticed your limits changing?
Bridgebench retweeted
I have not used Fable 5.1 once since Opus 5.5 dropped.
Opus 5.5 barely drains usage.
I ran it all week and my plan almost made it to reset.
Fable 5.1 used to kill a session in 30 minutes.
We're pledging $10,000 to track whether AI labs nerf their models after launch.
Here's how NerfBench works:
Every model gets benchmarked on launch day. That score becomes its 100%.
Then we keep retesting and measure power: performance, tokens, and cost per task.
If a model needs more tokens or costs more to do the same work, it's lost power. Even if the output looks the same.
90% to 110% is normal variance. Below 90% gets flagged.
Full breakdown: bridgebench.ai/blog/how-nerf…
Did Anthropic nerf Claude Opus 5.5?
The first NerfBench results are live.
We retested Opus 5.5 and GPT 6 Astra against their own launch scores.
Claude Opus 5.5: 99.2% (-0.8% vs launch)
GPT 6 Astra: 102.8% (+2.8% vs launch)
Verdict: No nerf detected.
Opus 5.5's small dip and GPT 6's small increase is within normal variance.
More models and more frequent retests are coming.
NerfBench only gets better as we collect more data.
Which models should we retest next?
Bridgebench retweeted
A lot of people are saying Anthropic has already nerfed Claude Opus 5.5.
We're launching NerfBench on BridgeBench tomorrow.
We have the day 1 results.
Tomorrow morning we show you the retest.
Is Claude Opus 5.5 nerfed or not?
Bridgebench retweeted
GPT 6 Astra made this video entirely with code.
Same prompt that I used for the Claude Opus 5.5 video.
Anthropic is so far ahead right now.
A lot of people are saying Anthropic has already nerfed Claude Opus 5.5.
We have every BridgeBench result for Opus 5.5 from day 1 of launch.
We're retesting soon.
If the scores drop, you'll see it here first.
Have you noticed a difference?
Bridgebench retweeted
A week ago you could hit your Claude Code limit in 30 minutes with Fable 5.1.
GPT 6 Astra in Codex could burn your whole week in 4 hours.
Today I ran Opus 5.5 for 8 hours and didn’t hit my limits a single time.
We are accelerating.
Bridgebench retweeted
Claude Opus 5.5 just ONE SHOT this marketing video.
Everything that you see was built entirely with code.
This is getting crazy....
Grok 4.7 is MORE EXPENSIVE than GPT 6 Astra.
Grok 4.7 may be the worst model release ever.
There is no reason that anyone should be using Grok.
Claude Opus 5.5 is the #1 front end model on BridgeBench.
1. Claude Opus 5.5: 950
2. Claude Fable 5.1: 860
2. Claude Opus 5: 860
4. Claude Fable 5: 800
5. Kimi K3: 780
Anthropic holds the top 4 spots. Not a single OpenAI model in the top 7.
Nobody is close to Claude on UI design right now.
Bridgebench retweeted
Claude Opus 5.5 completely changes what is possible.
Work that used to take me 2 weeks now takes me 2 hours. I still can't wrap my head around that.
It's Fable 5.1 level intelligence at $4 in and $20 out, it's faster than Opus 5, and the limits finally let me use it all day. No model has ever given me all three at once.
The world is changing so fast and most people have no idea yet.