ChessBench

@chessbench.ai

A new benchmark tracking how well language models play chess. Watch the games, follow the reasoning move by move, track the leaderboard.

Gemini 3.6 Flash is now the most accurate model we've tested to date! Its Accuracy score of 0.785 tops the leaderboard, ahead of Claude Fable 5 (0.748) -- it just doesn't always play legal moves... #2 overall debut; Gemini 3.5 Flash Lite lands at #13. @developers.google.com

Bild

Calling GPT-5.6 on the API? Check your bill closely. OpenAI is billing me for far more than the models actually generate. One example of many: on a hard position, GPT-5.6 used about 21k reasoning tokens. I was billed for 466k. Anyone else seeing this?

GPT 5.5: super coherent, almost never plays illegal moves, but least accurate of the three. Claude Fable 5: top-tier accuracy (~tied w/ Gemini), lowest coherence, hallucinates illegal moves unless I feed it the legal ones each turn. Gemini 3.1 Pro: high on both. And it's the oldest.

The part that has surprised me most so far: the three top models have completely different fingerprints. ChessBench tracks three things — coherence (how often does it play legal moves?), accuracy (how good are the legal moves it makes?), and Elo. The leaders diverge hard on the first two...