ChessBench
@chessbench.ai
A new benchmark tracking how well language models play chess. Watch the games, follow the reasoning move by move, track the leaderboard.
Gemini 3.6 Flash is now the most accurate model we've tested to date! Its Accuracy score of 0.785 tops the leaderboard, ahead of Claude Fable 5 (0.748) -- it just doesn't always play legal moves... #2 overall debut; Gemini 3.5 Flash Lite lands at #13. @developers.google.com
Calling GPT-5.6 on the API? Check your bill closely. OpenAI is billing me for far more than the models actually generate. One example of many: on a hard position, GPT-5.6 used about 21k reasoning tokens. I was billed for 466k. Anyone else seeing this?
Claude Sonnet 5 has been added to ChessBench, coming in at a surprising 14th overall...
Beyond the leaderboard, two things you can do on ChessBench: - watch any model's games (some include its reasoning embedded as move-by-move annotations) -- you can see exactly how and where it goes wrong - track every metric over time as new models ship chessbench.ai
ChessBench // A New Chess Benchmark for Language Models
How well can AI play chess? ChessBench benchmarks language models at chess, ranking them by Elo rating. Updated daily.
chessbench.ai
GPT 5.5: super coherent, almost never plays illegal moves, but least accurate of the three. Claude Fable 5: top-tier accuracy (~tied w/ Gemini), lowest coherence, hallucinates illegal moves unless I feed it the legal ones each turn. Gemini 3.1 Pro: high on both. And it's the oldest.
The part that has surprised me most so far: the three top models have completely different fingerprints. ChessBench tracks three things — coherence (how often does it play legal moves?), accuracy (how good are the legal moves it makes?), and Elo. The leaders diverge hard on the first two...
Most language model benchmarks are saturated. ChessBench isn't — and it's not close. GPT 5.5, Claude Fable 5, Gemini 3.1 Pro: not one of them beats a decent club player at chess yet. I built chessbench.ai looking for headroom. Turns out there's a lot.
ChessBench // A New Chess Benchmark for Language Models
How well can AI play chess? ChessBench benchmarks language models at chess, ranking them by Elo rating. Updated daily.
chessbench.ai