@chessbench

A new benchmark tracking how well language models play chess. Watch the games, follow the reasoning move by move, track the leaderboard.

Joined June 2026
Most language model benchmarks are saturated. ChessBench isn't — and it's not close. GPT 5.5, Claude Fable 5, Gemini 3.1 Pro: not one of them beats a decent club player at chess yet. I built ChessBench looking for headroom. Turns out there's a lot.
1
6
756
For those of you who are into supporting independent benchmarks: chessbench.ai/support Every dollar goes to API costs. More funding means more models tested, faster, with more games behind each rating.
2
1
7
398
Grok 4.6, 4.5, 4.3, and 4.2 have been added to ChessBench! More than any other family of models, @grok flounders when it's not given the legal moves for each position along the way... @elonmusk what's up
3
3
155
Gemini 3.7 Flash has been added to ChessBench, coming in with the highest Elo rating of any model we've benchmarked to date!
2
1
5
285
All three tiers of GPT-5.6 have (finally!) been added to ChessBench!
1
1
62
Gemini 3.6 Flash is now the most accurate model we've tested to date! Its Accuracy score of 0.785 tops the leaderboard, ahead of Claude Fable 5 (0.748) -- it just doesn't always play legal moves... #2 overall debut; Gemini 3.5 Flash Lite lands at #13.
🤖 Made with AI
1
1
106
Calling GPT-5.6 on the API? Check your bill closely. In ChessBench, every move is one API call, so the cost of each is easy to see. OpenAI is billing me for far more than the models actually generate. One example of many: on a hard position, GPT-5.6 used about 21k reasoning tokens. I was billed for 466k. I checked it against my OpenAI dashboard and the inflated numbers match to the cent, so it's what's charged, not a display bug. Another developer reproduced it. It also quietly hits o1, o3-mini, and o4-mini at 2x. Anyone else seeing this? @OpenAIDevs
2
102
Claude Sonnet 5 has been added to ChessBench, coming in at a surprising 14th overall...
1
4
5,119