Open any opening explorer and you get a column of moves with percentages next to them. It looks like a scoreboard, it reads like a recommendation, and it is neither. A database does not tell you which move is good. It tells you which moves were played, by people you know nothing about, in games whose results had almost nothing to do with the opening.
That is not an argument against databases. A database is the single most useful study tool most players never learn to operate. It is an argument for reading one properly, which takes about ten minutes to learn and saves years of misplaced confidence.
01What the numbers actually are
Every explorer shows you two different things and most people read them as one. Separate them deliberately:
- Frequency. How often this move was played in the games that reached this position. This is a fact about fashion, coverage and who happens to be in the database.
- Score. How the side to move did afterwards, usually as a percentage where a draw counts as a half point. This is a fact about the players, not the move.
The gap between them is where all the useful information lives. A move played in two percent of games that scores well is either an undervalued idea or a statistical accident. A move played in sixty percent of games that scores at the average is simply the main line, and tells you nothing except that it is the main line.
02Sample size, and the number below which you should stop looking
The single most common database error is trusting a percentage computed from a handful of games. A move with eleven games behind it can show any score at all, and the number will move by ten points if two more games arrive.
A workable rule of thumb, in the absence of anything more sophisticated:
| Games in the position | What the score means |
|---|---|
| Under 50 | Nothing. Read the moves, ignore the percentage entirely. |
| 50 to 300 | Weak signal. A twelve point gap might be real; a five point gap is noise. |
| 300 to 3,000 | Usable. Differences of eight points or more are worth investigating. |
| Over 3,000 | Reliable as a description of the past, which is still not a prediction. |
Thresholds are practical guidance rather than statistics. The direction is what matters: small samples lie loudly and confidently.
This matters most exactly where people look hardest. The deeper you walk into a line, the smaller the sample gets, so the percentages become least trustworthy at the precise moment they become most interesting. By move fifteen of a sharp variation you are usually reading the results of forty games, half of which were played by the same three people.
03Who is in your sample
A database is not a survey of chess. It is whatever collection of games somebody assembled, and the assembly rules leak into every number it produces.
Master databases contain games that were considered worth recording, which means strong players, serious events, and a strong bias toward openings that were fashionable when those events happened. Online databases contain everything, which is more representative of what you will face and much noisier. Neither is wrong. They answer different questions.
- Ask what you want to know firstIf the question is "what is objectively critical here", filter to strong players. If it is "what will actually appear on the board on Saturday", filter to your own rating band.
- Set the rating filter before you read anythingThe move order that works against 2400s and the move order that works against 1500s are often different moves, not different depths of the same move.
- Bound the yearsOpening theory moves. A line abandoned in 1998 may score beautifully in a database that still holds its glory days.
- Then read the columnNumbers read before the filters are set are worse than no numbers, because they arrive with confidence attached.
An unfiltered percentage is the average of a thousand games that have nothing to do with the game you are about to play.
04The survivorship problem
Here is the failure mode that catches strong players. A refuted line disappears from the database, and its score improves as it goes.
The mechanism is simple. Once a line is known to be bad, the people who know the refutation stop allowing it, and the only games left are the ones where somebody did not know. The remaining games are disproportionately won by whoever was better prepared, which is often the person playing the dubious line, because they chose it on purpose.
Take a sharp main line where preparation is everything.
Read a position like this from the frequency column alone and you learn what is fashionable. Read it from the score column alone and you learn which sidelines are currently catching people out. Only the two together, plus a look at the actual games, tell you anything about the position.
05Transpositions, or why your line has more games than you think
A database is organised by position, or it should be. If yours is organised by move order, it is lying to you by omission, because the same position is regularly reached through three or four different sequences and each sequence carries its own slice of the evidence.
This cuts both ways and both are useful:
- More evidence than you expected. The obscure line you are researching may have four hundred games under a different move order, with a different name, from a different opening family.
- Traps you did not sign up for. If your move order allows a transposition into a structure you have never studied, the database will show it long before an opponent does.
The practical habit: search by position rather than by opening name. Set the pieces up, ask what happened, and let the database gather every route into it. Opening names are a filing convention from the age of printed books. Positions are the real unit.
06The engine and the database answer different questions
People treat these as rivals. They are not even in the same category.
| Database | Engine | |
|---|---|---|
| Question it answers | What did people do here, and how did it go | What is the objectively best move |
| Good at | Practical value, surprise value, human difficulty | Truth, tactics, refutations |
| Blind to | Whether any of it was correct | Whether a human can hold the position |
| Use it for | Choosing what to play | Checking whether it survives |
The correct order is almost always database first, engine second. The database narrows the field to moves that humans actually play and shows you which of them create problems. The engine then tells you which of those candidates is sound. Reversing the order gives you a list of accurate moves with no information about whether anyone can play the resulting position.
07The five searches that do all the work
Most people use a database for exactly one thing, which is walking an opening forward from the starting position. That is the least interesting of its functions. These five cover almost everything worth asking.
- The position searchSet up a position and ask what happened. This is the primary search and the one that collects transpositions automatically. Use it whenever a game leaves your preparation, not only in the opening.
- The player searchEverything one person has played, sorted by opening and by result. This is opponent preparation, and it also works on yourself, which is uncomfortable and useful.
- The structure searchSet up the pawn skeleton you care about and ignore the pieces. Structures repeat across unrelated openings, and studying twenty games of the same structure teaches more than twenty games of the same opening.
- The filtered result searchGames in a line that finished decisively, or that lasted under thirty moves. Short decisive games are where the traps and standard mistakes live.
- The endgame searchFilter your opening down to the games that reached a specific endgame type. This is how you find out what your repertoire actually commits you to four hours in.
The third and the fifth are the ones almost nobody runs, and they are the two that change how you play. An opening is not a sequence of moves, it is a promise about the kind of position you will have to convert. The database is the only tool that can show you the promise before you sign it.
08What a database cannot tell you
Being clear about the limits is what makes the tool safe to use.
- Whether a move is sound. Only analysis answers that, and even then only within the horizon of the analysis.
- Why a game was lost. The result is attached to the whole game, not to the opening. A line does not become bad because three people who played it later blundered a rook.
- What your opponent knows. Their record shows what they have played, not what they have studied. A line they have never played may be the line they spent the summer on.
- What is about to become fashionable. By the time an idea has enough games to show up in the statistics, it is no longer a surprise.
- Whether you will enjoy the position. The one criterion that decides whether you will keep playing the line, and the one no database has an opinion about.
None of that is a defect. A database is a record of what people did, which is an enormous amount of information about human chess and no information at all about truth. Once you stop asking it the wrong questions, it stops giving you wrong answers.
09Reading a player, not a position
The other half of a database is the player index, and it answers a question the opening explorer cannot: not what is played here, but what does this specific person play.
Pull up a player and you get their whole record, sorted by opening. Three things fall out of it quickly:
- 01The repertoire. Usually narrower than they would like you to think. Most players have one first move and two defences, exactly as a well built repertoire recommends.
- 02The dislikes. Look for lines they have faced once and never allowed again. That avoidance is a stronger signal than any percentage.
- 03The drift. Compare the last year with the three before it. A recently changed defence is a defence they are still learning.
That is a different activity from opening study, and it deserves its own method, covered in how to prepare for an opponent.
10Five numbers that look meaningful and are not
- A score above 60 percent for either colour in a main line. Real main lines cluster near the overall average, which sits around 54 to 55 percent for White in large collections. A big outlier usually means a small or strange sample.
- A draw rate quoted without the rating band. Draw rates rise sharply with strength. The same line can look drawish and sharp depending entirely on who you included.
- Percentages on a move played fewer than fifty times. Covered above, and worth repeating because it is the mistake everybody makes twice.
- "Best by test" in a database that stops in 2015. Test conditions change. So do the tests.
- Your own score in your own games. You have maybe forty games in the line. That is a feeling, not a measurement. Use it to decide what you enjoy, not what is sound.
11A ten minute routine that works
- Set the position, not the namePlay the moves onto the board or paste the FEN. Let transpositions collect themselves.
- Set the filtersRating band you actually play in, last ten to fifteen years, and the result and event filters only if you have a specific question.
- Read frequency and score separatelyOne column tells you what to expect. The other tells you which expectations were expensive.
- Check the sample before you believe anythingIf it is under fifty games, treat the percentages as decoration.
- Open three actual gamesThis is the step everyone skips and the one that teaches. Watch what the structure turns into by move twenty five.
- Now turn on the engineVerify the two or three candidates you liked. Not before, or you will only ever find moves you cannot play.
Six steps, ten minutes, and the output is not a percentage. It is a candidate move, a plan, and a rough idea of what your opponent is likely to do. That is what a database is for.
Questions
What is a chess database?
A searchable collection of recorded games, indexed so you can look up a position and see every game that reached it, who played it, what was played next and how those games finished. Modern databases hold millions of games and run in a browser, with no download required.
How do I search a chess database by position?
Set the position up on the board, either by playing the moves or by pasting a FEN, then ask the database for games that reached it. A position based search collects every move order that transposes, which a search by opening name will miss.
Do chess opening statistics mean a move is good?
No. A score percentage describes how the players who chose that move performed afterwards, which is mostly a fact about those players and their preparation. Use the statistics to see what is common and what causes practical problems, then use an engine to check whether your candidate is sound.
How many games do I need before a percentage is meaningful?
Treat anything under fifty games as unreadable, and anything under three hundred as a weak signal where only large gaps matter. Sample sizes shrink fast as you go deeper into a line, exactly where the numbers look most interesting.
Should I filter a chess database by rating?
Yes, and before you read any numbers. The moves that work against 2400 rated players and the moves that work against 1500 rated players are frequently different, so an unfiltered percentage averages together two populations that have little to do with each other, or with you.