9,762 rated games analysed with Stockfish 17, plus a 33,598-game comparison set.
The finding
The game-losing mistake arrives at almost exactly the same relative point in the game at every rating: about 78–83% of the way through.
What changes with rating is not when players go wrong within a game. It is how far the game gets before someone does.
| Rating | Median game length (plies) | Reach an endgame | Decided in the endgame, of those that reach one | Relative position of the mistake |
|---|---|---|---|---|
| 700-899 | 49 | 35% | 45% [41%–50%] | 0.781 |
| 900-1099 | 54 | 39% | 46% [42%–50%] | 0.818 |
| 1100-1299 | 55 | 42% | 49% [45%–53%] | 0.816 |
| 1300-1499 | 59 | 48% | 47% [43%–50%] | 0.816 |
| 1500-1799 | 62 | 53% | 53% [50%–56%] | 0.826 |
| 1800-2000 | 66 | 59% | 53% [50%–56%] | 0.831 |
Read the last column first. It barely moves: 0.78 to 0.83, five percentage points of a game. A 900 and an 1800 both lose the game roughly four-fifths of the way through it.
The third column is where the real difference sits. 35% of sub-900 games reach an endgame at all; 59% do at 1800–2000. Weaker players' games end before the endgame arrives — which is exactly why raw counts of "games decided in the endgame" rise with rating, and why that number on its own is misleading.
Among games that do reach an endgame, the share decided there rises only from 45% to 53%. That comparison is free of the length problem: every game in it had the opportunity.
Phase of the decisive mistake
Classified by material rather than move number — endgame means six or fewer non-king, non-pawn pieces, which is the threshold Lichess itself uses. 95% confidence intervals in grey.
| Rating | Games | Median move | Relative position | Opening | Middlegame | Endgame |
|---|---|---|---|---|---|---|
| 700-899 | 1,526 | 17 | 0.78 | 24% | 50% | 26% [24%–28%] |
| 900-1099 | 1,577 | 19 | 0.82 | 17% | 53% | 30% [28%–32%] |
| 1100-1299 | 1,639 | 20 | 0.82 | 17% | 52% | 32% [30%–34%] |
| 1300-1499 | 1,670 | 22 | 0.82 | 13% | 53% | 33% [31%–36%] |
| 1500-1799 | 1,660 | 23 | 0.83 | 11% | 52% | 37% [35%–39%] |
| 1800-2000 | 1,690 | 24 | 0.83 | 11% | 53% | 37% [35%–39%] |
The trend across the six bands is consistent in direction and the extreme-band contrasts are comfortably outside their confidence intervals. Individual adjacent steps are not: several differ by less than their intervals and should not be read as real.
Almost nothing is decided in a theoretical endgame
GM Noël Studer, reviewing a draft, predicted that the share of games decided in the kind of endgame position an improver actually studies would be “extremely minimal”. That prediction is correct, and by a wide margin.
| Rating | Median pieces on the board at the mistake | Decided with ≤7 pieces (tablebase range) | Decided with ≤10 pieces | “Endgame” by the Lichess threshold |
|---|---|---|---|---|
| 700-899 | 24 | 3.5% | 6.3% | 26.1% |
| 900-1099 | 23 | 2.8% | 6.7% | 29.8% |
| 1100-1299 | 23 | 2.8% | 7.6% | 31.9% |
| 1300-1499 | 22 | 2.2% | 5.6% | 33.4% |
| 1500-1799 | 22 | 3.4% | 7.3% | 37.0% |
| 1800-2000 | 21 | 2.4% | 6.8% | 36.8% |
Under 4% of games at any rating are decided with seven or fewer pieces on the board — the range that endgame theory actually covers. The median decided position has 24 pieces on it at beginner level and 21 at 1800–2000.
This is the gap between the word “endgame” as a database threshold and as a chess player means it. A position with queens and both rooks still on counts as an endgame under the Lichess definition, and is a middlegame to anyone playing it.
Time control changes the answer
Endgame share of decisive mistakes, split by time control. Pooling time controls hides a real effect, particularly at beginner level.
| Rating | Bullet | Blitz | Rapid |
|---|---|---|---|
| 700-899 | 15% n=320 | 27% n=760 | 33% n=443 |
| 900-1099 | 27% n=426 | 31% n=770 | 30% n=375 |
| 1100-1299 | 32% n=455 | 34% n=834 | 28% n=343 |
| 1300-1499 | 31% n=560 | 34% n=855 | 37% n=251 |
| 1500-1799 | 36% n=584 | 37% n=841 | 40% n=227 |
| 1800-2000 | 37% n=701 | 37% n=779 | 37% n=205 |
At beginner level the difference is large: a sub-900 bullet game is decided in an endgame 15% of the time against 33% in rapid. Bullet games end before the endgame arrives.
What this looks like on the board
Two real games from the sample, caught at the moment before the decisive mistake.
A game decided early (700–899)

Most pieces still on their starting squares. Black to play, move 12. Black played hxg6 (red square to green), costing 6.11 pawns — Stockfish wanted Nc4. White won. Real game from the sample: lichess.org/BtMxLoDl (691 vs 711).
A game decided late (1800–2000)

Thirty-eight moves in, material level. White to play, move 38. White played Rxg6 (red square to green), costing 4 pawns — Stockfish wanted Rh3. Black won. Real game from the sample: lichess.org/NMi46FUJ (1812 vs 1787).
Method
Source. The Lichess public archive for March 2025 (lichess_db_standard_rated_2025-03), rated standard games.
Sampling. 2,000 games drawn at random from each of six rating bands, banded by the average of the two players' ratings. Bands are filled independently, so the sample is balanced by rating rather than reflecting Lichess's own distribution.
Evaluation. Every position analysed with Stockfish 17, depth 12 (sf17 population), White-relative, one evaluation per ply.
The decisive mistake. Last move after which the player was at a 2-pawn or worse deficit and never recovered.
Phase. by material: endgame = 6 or fewer non-king non-pawn pieces (Lichess threshold); opening = 13+ such pieces; middlegame otherwise. Move-number boundaries are reported in the data file for comparison only.
Exclusions. 1,451 games shorter than 20 plies; 3,666 that never reached a two-pawn decision; 681 where the deficit came from the opponent's strong move rather than the loser's error.
Confidence intervals. Wilson 95%, on phase shares.
What this does not show
- Nothing about what you should study. This measures where mistakes occur, not the return on study time. Those are different questions and only the first one is answered here.
- The comparison set is not an independent replication. The 33,598-game cloud-eval set uses the same engine family and the same construction, so it tests sampling, not measurement. It is also self-selected — those evaluations exist because a user asked for analysis.
- Depth 12 is shallow, and evaluation noise at that depth can move which move is identified in a close case.
- One month, one site. March 2025 on Lichess. A different pool may differ.
- Bands use the average of both players' ratings. A 1700 playing a 1300 lands in 1500–1799.
- Raw phase counts are misleading on their own, which is why the normalised figures lead this article. Any comparison of "share decided in the endgame" across ratings has to account for the fact that weaker players' games often end before an endgame exists.
- No causal claim.
Acknowledgements
Thanks to GM Noël Studer (Next Level Chess) for the methodology: the definition of the game-deciding mistake used here is his, as is the point that a material-based endgame threshold still admits positions most players would call middlegames — which prompted the theoretical-endgame test above.
Data
Released under CC BY 4.0 — reuse freely, including commercially, with attribution.
- decisive-v2.json — current results, all bands, both phase definitions, by time control
- truncation-test.json — the game-length normalisation
- studer-endgame-test.json — theoretical-endgame test
Method questions: [email protected].