← All articles
Testing & data

Why your 5-game win rate is lying to you

· 5 min read

"I'm 4–1 with the new list" is the most dangerous sentence in deck testing. It feels like evidence, and it's mostly noise. Rates computed from a handful of games swing wildly for no reason at all, and acting on them is how good lists get butchered and bad ones get taken to tournaments.

The coin that looks rigged

Flip a fair coin five times. About one run in five shows four or more heads. Nobody looks at that and calls the coin rigged — five flips is obviously too few. But a player who goes 4–1 in testing will confidently call a 50% matchup favorable, and a player who opens 1–4 will cut the deck entirely. Both are making exactly the mistake the coin flip was supposed to rule out.

How many games is enough?

There's no single magic number — it depends on how narrow a slice of your games the stat is measuring. The narrower the slice, the more games it takes to mean anything. Reasonable working thresholds:

  • ~20 games for a deck-level read (overall win rate, setup rate).
  • ~30 games for a per-card rate — utilization numbers only count games where that card appeared.
  • ~50 games for a single matchup, because any one matchup is a small fraction of your games.

These are thresholds for taking a number seriously, not for having an opinion. You'll form impressions earlier — just label them impressions.

Honesty rules for reading your own data

  • Label thin data instead of hiding it. A win rate over 12 games is worth showing — with a "small sample" flag on it, not rounded into false confidence.
  • A zero-game rate is "no data yet," not 0%. The difference matters enormously and most spreadsheets get it wrong.
  • Exclude what you can't verify. A game you can't reconstruct honestly is better dropped than guessed at.
  • An early read is an early read. Say so, even to yourself.

What this means for your process

Structure testing so decisions wait for the sample they need. List surgery on a 5-game impression is almost always premature; so is declaring a matchup unwinnable off three losses. In the meantime, lean on the things that don't need samples at all: draw odds, mulligan rates, and the theory floor are exact from the decklist alone — math doesn't need 50 games. Your logged games catch up eventually; let the two meet before you operate.

See these numbers on your own deck.

DeckSequence runs this exact math on your list and checks it against your logged games. Free, in private beta.

Request early access