Sstatmate Analyse my games →
Comparison

Comparing chess players and games fairly: cross-platform ratings and true game quality

Two of the most interesting comparisons in chess are also the two that raw numbers get badly wrong: comparing ratings across different sites, and judging which of two games was actually the better performance. Here is why they mislead, and how to do them properly.

Comparison is one of the most useful things you can do with your chess data. It turns a vague sense of “am I any good” into concrete answers. But the two most compelling comparisons both have a hidden catch that makes the obvious approach flat-out wrong. The first is comparing a player on one site to a player on another. The second is deciding which of two games was better played. In both cases the raw number, a rating or an accuracy percentage, lies to you, and for the same underlying reason: the number only means something next to its context.

Why a 1500 on Lichess is not a 1500 on Chess.com

A rating is only meaningful inside its own pool. Each site has its own player population and its own rating maths, so the numbers were never on the same scale to begin with. In practice, Lichess ratings tend to run noticeably higher than Chess.com ratings for the same real strength, often by a hundred points or more, and the size of that gap is not even constant. It changes from one time control to the next.

This is why the moment you try to settle “who is actually stronger” between a friend on Lichess and yourself on Chess.com, the raw ratings deceive you. A Lichess 1800 and a Chess.com 1800 are simply not the same player, and whichever number is bigger tells you nothing on its own.

the same strength, two different numbersLichess1400160018002000Chess.com125014501650185017501600~150 pt gap
The same real playing strength lands on two different numbers. A player who is a Chess.com 1600 is roughly a Lichess 1750, because the Lichess scale sits higher. Comparing the raw numbers without accounting for that gap makes the Lichess player look stronger than they are.

The fix is normalisation: converting both players onto a single scale before you compare anything, using the typical offset between the two systems for each time control. Only once both are expressed in the same currency does “who is stronger” have an honest answer. It is the difference between comparing two prices after converting them to the same currency, versus just noting that one number is bigger.

Done right, this settles the friendly argument between players on different sites, and it lets you benchmark yourself against someone whose games you admire even when they happen to play on the platform you do not use.

Why accuracy alone cannot tell you which game was better

The second trap is subtler and catches even experienced players. Accuracy feels like the obvious way to rank two of your games: higher accuracy means the better game. But accuracy is heavily distorted by one thing that has nothing to do with how well you played, which is how sharp the position was.

In a quiet, simple game, most moves are natural and hard to get wrong, so your accuracy runs high even for fairly ordinary play. In a sharp, tactical game full of forcing sequences and only-moves, every single move is a genuine test, and even excellent play produces a lower accuracy number. The consequence is uncomfortable: a hard-fought, brilliant tactical win can score a lower accuracy than a sleepy, low-effort positional game. Ranking the two by accuracy alone punishes the harder and more impressive performance.

To compare two games fairly you have to put the difficulty back in. You can measure difficulty by how much “only-move pressure” a game contained, meaning how often there was a single correct move and everything else lost ground. A game packed with only-moves was genuinely harder to play well, so a slightly lower accuracy in that game can represent stronger chess than a high accuracy in an easy one. A fair game-quality comparison shows you both numbers side by side, the accuracy and the difficulty, so you can tell whether a lower score means worse play or simply a harder game.

That unlocks two questions worth asking. Was my flashy tactical win actually better than my quiet grind, or did it just feel that way? And, most usefully, am I improving? Put a game from six months ago next to a recent one and compare the blunders, the mistakes and the clean conversions, adjusted for how hard each game really was.

In your own games

Both comparisons, done for you in Statmate

Statmate’s Compare tab has both of these built in. Cross-platform compare pits a Chess.com player against a Lichess player and normalises their ratings onto one scale, and you can flip between viewing everything on the Chess.com scale or the Lichess scale, with a full stat-by-stat breakdown underneath.

Statmate cross-platform compare with ratings normalised onto one scale and a stat-by-stat breakdown
Cross-platform compare: a Chess.com player against a Lichess player, with both ratings normalised onto one scale (toggle between the Chess.com and Lichess scale) and a full stat-by-stat breakdown below.

Two-game quality compare grades two games move by move and reports each game’s accuracy right alongside its difficulty, the only-move pressure, so a sharp game is never unfairly punished against a quiet one. It works on your own two games, which is how you benchmark your current self against your past self, or on two different players’ games.

Statmate two-game quality compare showing accuracy alongside difficulty for each game
Two-game quality compare: each game graded move by move, showing accuracy alongside its difficulty (only-move pressure), so a sharp tactical game is not unfairly punished against a quiet one.

Both run on Stockfish 18, from the free in-browser Lite engine up to the full engine on premium, so the grades are real engine analysis rather than a rough estimate.

The common thread

Both comparisons come back to the same idea. A raw number, whether it is a rating or an accuracy percentage, is only meaningful next to its context. Strip the context away and you get a confident answer that happens to be wrong. Put the context back, the platform’s scale, or the difficulty of the position, and comparison becomes one of the sharpest tools you have for actually understanding your chess.

Advertisement
Advertisement