@bridgebench

The official AI benchmark of the vibe coding movement @bridgemindai

United States
Joined March 2026
Claude Sonnet 5.5 has Opus 5.5 level taste at less than half the price. Same lava lamp prompt: Claude Opus 5.5: $1.27, 9m 35s Claude Sonnet 5.5: $0.47, 5m 44s Sonnet nailed the glow, the lighting, and the details. And finished 4 minutes faster. Great UI doesn't have to be expensive anymore.
31
13
253
14,396
Bridgebench retweeted
Claude Sonnet 5.5 created this video entirely with code. Sonnet 5.5 Max effort is better than Claude Opus 5.5.
34
30
2
739
28,797
Bridgebench retweeted
Claude Sonnet 5.5 just ONE SHOT Mario Kart. This is by far the best result that I have seen on this test.
42
29
9
489
27,161
Bridgebench retweeted
Claude Sonnet 5.5 just ONE SHOT this zombies game. Max effort is noticeably better than every other level but it is expensive. This game cost $177 and took 49 minutes to build. This is insane.
59
29
9
865
43,596
Claude Sonnet 5.5 just dropped. We put it through the BridgeBench black hole merger test against Opus 5.5 and Fable 5.1. Claude Fable 5.1: $1.22, 4m 44s Claude Opus 5.5: $0.69, 6m 8s Claude Sonnet 5.5: $0.34, 4m 14s Sonnet 5.5 was the fastest and cheapest of the three. Half the cost of Opus 5.5. Which one did the best?
56
9
5
544
53,472
Claude Sonnet 5.5 beat Fable 5.1 on Artificial Analysis. 1/5 the price. We are accelerating.
17
12
5
479
13,904
Bridgebench retweeted
Claude Sonnet 5.5 is imminent. Sonnet 5 was a disaster. Terrible tokenizer, wasted tokens, F tier in my rankings. I never used it. But I am hearing that Sonnet 5.5 will be the same leap that Opus 5 had with Opus 5.5. Sonnet 5.5 is going to be a leap.
33
11
1
522
37,876
A lot of people think Claude and Codex are silently cutting their usage limits. Right now there's no way to prove it. So we're building UsageBench on BridgeBench. We track the 5 hour and weekly limits over time. If a lab ever cuts your limits in the background, you'll see it here first. Have you noticed your limits changing?
121
63
12
1,710
50,479
Bridgebench retweeted
I have not used Fable 5.1 once since Opus 5.5 dropped. Opus 5.5 barely drains usage. I ran it all week and my plan almost made it to reset. Fable 5.1 used to kill a session in 30 minutes.
75
19
5
996
37,880
We're pledging $10,000 to track whether AI labs nerf their models after launch. Here's how NerfBench works: Every model gets benchmarked on launch day. That score becomes its 100%. Then we keep retesting and measure power: performance, tokens, and cost per task. If a model needs more tokens or costs more to do the same work, it's lost power. Even if the output looks the same. 90% to 110% is normal variance. Below 90% gets flagged. Full breakdown: bridgebench.ai/blog/how-nerf…
Did Anthropic nerf Claude Opus 5.5? The first NerfBench results are live. We retested Opus 5.5 and GPT 6 Astra against their own launch scores. Claude Opus 5.5: 99.2% (-0.8% vs launch) GPT 6 Astra: 102.8% (+2.8% vs launch) Verdict: No nerf detected. Opus 5.5's small dip and GPT 6's small increase is within normal variance. More models and more frequent retests are coming. NerfBench only gets better as we collect more data. Which models should we retest next?
39
19
5
640
40,707
Bridgebench retweeted
A lot of people are saying Anthropic has already nerfed Claude Opus 5.5. We're launching NerfBench on BridgeBench tomorrow. We have the day 1 results. Tomorrow morning we show you the retest. Is Claude Opus 5.5 nerfed or not?
338
226
109
4,619
820,830
Bridgebench retweeted
GPT 6 Astra made this video entirely with code. Same prompt that I used for the Claude Opus 5.5 video. Anthropic is so far ahead right now.
Claude Opus 5.5 made this video entirely with code.
45
16
4
420
62,810
A lot of people are saying Anthropic has already nerfed Claude Opus 5.5. We have every BridgeBench result for Opus 5.5 from day 1 of launch. We're retesting soon. If the scores drop, you'll see it here first. Have you noticed a difference?
274
93
21
4,928
307,595
Bridgebench retweeted
A week ago you could hit your Claude Code limit in 30 minutes with Fable 5.1. GPT 6 Astra in Codex could burn your whole week in 4 hours. Today I ran Opus 5.5 for 8 hours and didn’t hit my limits a single time. We are accelerating.
90
42
5
2,034
64,317
Claude Opus 5.5 just ONE SHOT this marketing video. Everything that you see was built entirely with code. This is getting crazy....
32
5
1
270
16,907
Bridgebench retweeted
Claude Opus 5.5 made this video entirely with code.
105
76
24
1,470
128,656
Grok 4.7 is MORE EXPENSIVE than GPT 6 Astra. Grok 4.7 may be the worst model release ever. There is no reason that anyone should be using Grok.
77
31
11
872
47,541
GPT 6 Luna is absurdly good for the price.
How did OpenAI make GPT 6 Luna this good? GPT 6 Luna: under a penny Claude Fable 5.1: $1.67 Claude Opus 5.5: $0.90 This is the cheapest model in the world and it's good.
6
1
104
20,246
Claude Opus 5.5 is the #1 front end model on BridgeBench. 1. Claude Opus 5.5: 950 2. Claude Fable 5.1: 860 2. Claude Opus 5: 860 4. Claude Fable 5: 800 5. Kimi K3: 780 Anthropic holds the top 4 spots. Not a single OpenAI model in the top 7. Nobody is close to Claude on UI design right now.
18
8
1
277
14,101
Claude Opus 5.5 completely changes what is possible. Work that used to take me 2 weeks now takes me 2 hours. I still can't wrap my head around that. It's Fable 5.1 level intelligence at $4 in and $20 out, it's faster than Opus 5, and the limits finally let me use it all day. No model has ever given me all three at once. The world is changing so fast and most people have no idea yet.
52
36
5
936
29,066