← All articles
Testing & data

Theory vs practice: testing a deck without fooling yourself

· 6 min read

Deck testing usually goes one of two ways. Either the pilot falls in love — every win confirms the deck, every loss was variance — or the pilot panics, cutting cards after three bad games. Both are the same mistake: judging results against no prediction at all. Put two numbers side by side instead: what the math says the list should do on draws alone, and what your logged games actually show.

The theory floor

Start with a question the math can answer exactly, like: "how often is an attacker online by turn 2?" Count the copies in your attacking line — say 9 across the evolution stages — and compute the odds of having seen what you need by then. This calculation counts only your opening hand plus the guaranteed draw for turn — no Ultra Ball, no Professor's Research, no search engine at all. That makes it a floor, not a forecast: a real deck, played with its engine, should sit above this number, and beating the floor is expected, not suspicious.

The band: theory is a range, not a point

Because the going-first player skips a draw, the floor is really two numbers — one per seat. For the 9-copy attacking line above, "attacker online by turn 2" computes to 75% on the play and 79% on the draw. That 75–79% spread is the band. Nothing is invented to make it look like a range; it's the real asymmetry of the game. Your practice results should be compared against the band, not against a single averaged number.

Reading your real results against it

Log your games, measure the same stat, and read the comparison:

  • Above the band — normal for a tuned deck. Your draw and search engine is doing its job; that's what those slots are for.
  • Inside the band — you're performing in line with raw draws. Fine, but ask whether the engine cards are earning their slots (utilization data answers this).
  • Below the band — worth a look. Something is eating the equity your list mathematically has.

A gap is a hypothesis, never a verdict

Falling below the band does not mean you're misplaying. It means one of three things: variance (small samples swing wildly), the list (maybe the engine can't actually find the pieces), or play (maybe hands are being sequenced in an order that wastes them). The comparison tells you where to look, not what the answer is. Any tool or teammate that calls it misplay after twelve games is guessing, not measuring.

The habit

Before testing a change, write down what the math predicts. After testing, compare against the prediction — by seat, with the sample size in view. The math tells you what should happen; your games tell you what did. The difference is where you get better.

See these numbers on your own deck.

DeckSequence runs this exact math on your list and checks it against your logged games. Free, in private beta.

Request early access