How these numbers are made

A player page makes two kinds of claim.

Claim 1

Someone differs from a reference group.

Depends on which group.

Part 1 · Who you are compared against →

Claim 2

The difference is big enough to believe.

Depends on how many hands are behind it.

Part 2 · Whether a difference is real →

Both are worked out below, with the arithmetic shown.

Part 1

Who you are being compared against

Three reference groups appear on a player page and they are not interchangeable. Two of them can differ by four points on the same statistic in the same game — which reads as a contradiction until you know what each one counts.

GroupWho is in itUsed onWhy
Poolevery player in this game with 500+ hands.One vote each.Distance from poolEverything flagged unusualwhy unweighted →
FieldThe same people.Weighted to the winners’ volume profile.Shape of a winnerwhy re-weighted →
WinnersThe profitable subset.Measured on hands not used to pick them.Shape of a winnerwhy split in half →

1.1 The player pool — what everybody does

The question it answers

What does a typical player in this exact game do?

It is the reference group behind the Distance from pool panel and behind everything that panel flags.

It is built in five steps.

  1. 1

    Cut the data to one game.

    Same currency, same stake, same table size — and the same room, wherever that room has enough players to stand on its own.

    NL2 and NL200 are not the same game and neither are two rooms: measured, pool VPIP at NL2 6-max differs by 6.3 points between rooms.

  2. 2

    Drop anyone with under 500 hands in that game.

    A player seen for twelve hands has a VPIP of 0% or 50% or 100% and none of those describe them.

  3. 3

    Work out each surviving player’s own percentage.

    From their counters — actions divided by opportunities.

  4. 4

    Take the median of those percentages.

    One number per player, then the middle one.

  5. 5

    Publish only if 30 players qualified.

    Below that the “typical player” is a handful of people.

The whole of it, in two lines:

each player = 100 × actions / opportunities (needs ≥ 500 hands) pool = median( player₁ , player₂ , … , playerₖ ) (needs k ≥ 30)

Why the median, and not the average

The average asks

What happens if you add everyone up?

The median asks

Who is in the middle?

Only the second is a description of the opponent you are likely to be sitting with, and poker populations always have a long loose tail to drag an average around.

Nine players at one stake, sorted by how often they enter a pot:

19.2 21.8 23.1 24.9 [26.8] 28.4 31.0 35.6 48.9 ↑ middle value = the pool median 26.8 ← what a typical player does average 28.9 ← dragged up 2.1 points by one loose player

Illustrative figures, chosen to make the arithmetic followable — not a measurement.

Why one number per player, and not one big pile of hands

Lumping every hand together and dividing once would let a single grinder with 200,000 hands outvote three hundred casual players.

COUNTED BY HANDone grinder200,000 hands300 casual players150,000 handsCOUNTED BY PLAYERone grinder1 vote300 casual players300 votes
Three hundred casual players at the 500-hand minimum this section already imposes. Counted by hand the grinder outweighs all of them — and only just.

Each player gets one vote regardless of volume, because the question is what a typical person does, not what a typical hand looks like.

That choice is exactly what the next section has to undo, on purpose. The field →

1.2 The field — the pool, reshaped to be a fair comparison

The question it answers

What does a typical player do, once volume is held level?

The field is the reference group in Shape of a winner, on the radar and in the readout under it. It is the same people as the pool, counted differently, and it exists to stop the site claiming something it cannot support.

The problem it solves

Winners play roughly twice the volume of a median player. High-volume players also behave differently from everyone else — tighter, more aggressive, fewer limps — whether or not they win. So “winners vs the pool” mixes two effects together, and mostly measures the wrong one.

Measured on a real segment: comparing winners to the raw pool reported them folding to 3-bets 8.4 points less than everyone else. Holding volume level, the real figure is 4.0 points. More than half of the “winning” effect was a regular-versus-everyone effect wearing its clothes.

WINNERS' EDGE IN FOLD VS 3BETagainst everyone at the stake8.4 ptsagainst a volume-matched field4.0 ptsmore than half was a regular-vs-everyone effect0510gap between winners and the field, in percentage points
The winners are the same people in both bars. Only the group they are measured against changes.

How the reshaping works

  1. 1

    Sort every qualifying player into a volume bucket.

    Doubling bands, so 1,024–2,047 hands is one bucket and 2,048–4,095 the next. Volume spans three orders of magnitude, and doubling bands keep the busy tail from collapsing into one lump.

    bucket(player) = floor( log₂( hands ) )
  2. 2

    Give each player their bucket’s winner share.

    Not their own result — the share of that bucket which is made of winners.

    weight(player) = winners in their bucket / players in their bucket
  3. 3

    Take the weighted median.

    The same median as the pool, with each player counting as much as their weight rather than once.

    field = weighted median of every qualifying player’s value

Why that weight

Every bucket contributes exactly its own winner count.

n players each carrying W/n add up to W, so the weight landing on each bucket is the winners’ volume distribution, by construction.

The population, counted once each1,000 players60%30%10%The population, counted by weight200 in total30%45%25%The winners, counted once each200 winners30%45%25%500 – 2k hands2k – 8k8k+
The lower two bars are identical, which is the point: the weights are built so the population takes the winners' shape.

The field is the whole population, wearing the winners’ volume profile.

Worked through

A thousand qualifying players in three volume buckets. Everyone in a bucket shares a VPIP, to keep the sums visible:

BucketPlayersWinnersVPIPWeight eachBucket weight
500 – 2k hands6006028.060/600 = 0.1060
2k – 8k hands3009024.090/300 = 0.3090
8k+ hands1005021.050/100 = 0.5050

Illustrative figures, chosen to make the arithmetic followable — not a measurement.

POOL — every player counts once, 1,000 votes sorted: 100 players at 21.0 | 300 at 24.0 | 600 at 28.0 vote 500 of 1,000 lands in the last block pool = 28.0 FIELD — each player counts as much as their weight, 200 in total sorted: weight 50 at 21.0 | 90 at 24.0 | 60 at 28.0 running: 50 140 200 halfway = 100, which lands in the middle block field = 24.0

Same thousand people, same thousand values. Counting each of them once gives 28.0; counting them in the winners’ volume proportions gives 24.0. Neither is wrong — they answer different questions, and putting the second beside a winner median is the only one of the two that is a like-for-like comparison.

1.3 Proven winners — and how they are proven

The question it answers

What do the players who actually make money do differently?

A winner is a player whose EV-adjusted win rate is positive in that game. Two details do the real work.

1. Luck is removed where it can be

Selection uses EV-adjusted winnings, not actual ones. If a player gets all-in with aces against kings and loses, the actual result says they lost the pot; the EV-adjusted result gives them the ~81% of it the cards said was theirs. All-in variance is the one chunk of luck that can be measured exactly, so it is not allowed to decide who counts as a winner.

2. Nobody is measured on the hands that got them picked

Each player’s hands are split in two by a hash of the hand id — effectively a coin flip per hand, independent of time, table, session and stake. Winners are chosen on one half and measured on the other.

a player's 4,000 hands at NL25 even hands (2,000) → EV-adjusted +38.4 bb → qualifies as a winner odd hands (2,000) → VPIP 23.1, PFR 18.4 → these go into the median the half that decided WHETHER they count is never the half that decides WHAT they look like

Illustrative figures, chosen to make the arithmetic followable — not a measurement.

Without the split the panel flatters itself. Picking winners by money won and then reporting statistics made of money won counts the same good luck twice:

they are labelled a winnerbecause they won showdownsthey show a high “won money at showdown”because they won showdowns

one run of luck, doing both jobs

Measured. Selecting and measuring on the same hands put winners 2.9 points above the field on won-money-at-showdown. Split across separate halves, the gap collapsed to 0.4 — about seven-eighths of it was never skill. One statistic, won-when-saw-flop, changed direction outright.

WINNERS' EDGE AT SHOWDOWN — W$SDpicked and measured on the same hands2.9 ptspicked on one half, measured on the other0.4 pts7/8 of the edge was the luck that picked them024gap between winners and the field, in percentage points
Both bars are the same statistic on the same players. Only the hands used to pick them differ.

3. What “winning” is scoped to

  • Per stake and room, not per situation.“Wins at NL50” is a property worth selecting on;“wins inside limped pots” is noise wearing the same word.Winner status is decided once for the game and applied to every pot type within it.
  • 500 measured hands minimum — so a published winner row rests on the same evidence a pool row does.
  • 30 winners minimum before a segment publishes at all.

1.4 Pool and field will not match

The gap between them is not an error. On a real NL25 full-ring segment, the same people counted two ways — with the casual, low-volume slice counted the way the winners’ own volume profile counts it:

poolfieldVPIP4.1 points apartPFR0.2 points apart0%10%20%30%
Same people in both, counted two ways. One number moves a long way and the other barely moves — which is the fingerprint.

That contrast is the fingerprint of the difference:regulars limp and cold-call far less, and raise about as often.

1.5 Why each panel uses the one it does

One rule decides it:

Matching removes a difference between two groups. Only remove the difference you are NOT trying to measure.

Distance from pool → the pool, unweighted

This panel asks a descriptive question: does this player differ from the people you will actually be sitting with?

Your next opponent is drawn from the population as it isa random person, not a random hand.

So the yardstick is one vote per player, with nothing weighted away — for two reasons.

  1. 1

    It keeps the volume effect.

    If a 200,000-hand regular folds to 3-bets less than the pool, part of that gap is just that regulars differ from casual players. It is still true and still worth acting on — you should 3-bet that player less than you would a stranger. Matching it away would delete part of the answer.

  2. 2

    And it keeps the unit intact.

    Being flagged unusual takes a distance measured in how much that statistic varies between players — interquartile ranges of the pool. That spread is a property of the same one-vote-per-player distribution the median comes from. Re-weight the population and you have changed the unit the distance is measured in, not merely the point it is measured from.

Shape of a winner → the field, volume-matched

This panel makes a far stronger kind of claim.

Not “these two groups differ” but“this difference is what winning looks like”— and the reader is expected to change how they play because of it.

  1. 1

    Winners are not a random slice of the pool.

    They were selected, and selection correlates with volume: winners play about twice the hands of a median player. So winners and the pool differ in volume by construction, before a single statistic is compared.

  2. 2

    So any gap carries two explanations at once.

    a gap between winners and the pool“this is whatwinning looks like”“this is whatplaying a lot looks like”the panel may claim thisand it may not claim this
    Both readings of the gap are true. Only one of them is about winning, and only that one is the panel's to publish.
  3. 3

    So the volume effect is matched away.

    Here it is a rival explanation rather than part of the answer. The same effect is signal in one panel and noise in the other, because the two panels are claiming different things.

Part 2

Whether a difference is real

Every number on a player page is measured off a finite number of hands, so every number carries an error. This is how we work out how big that error is, and when a difference is small enough that we refuse to call it a difference at all.

2.1 The question

Suppose a player c-bets the flop 38.7% where the field c-bets 60.3%. That is a gap of 21.6 points and it looks enormous.

Whether it means anything depends entirely on how many flops they actually c-bet into.

off 40,000 opportunities

21.6 points

a real, exploitable trait

off 12 opportunities

21.6 points

a coin landing the same way a few times

So every gap needs a companion number: how far the gap would move on chance alone. Below that, we grey it out.

2.2 The experiment

Population

40 PokerStars accounts, 100,000+ hands each, one stake and format apiece

Spread across stakes rather than taking the 40 biggest — the largest accounts cluster in a couple of games, and a rule fitted to one game would not transfer.

Blocks

Consecutive runs of 100, 200, 500 … 20,000, in the order played

A filter selects a run of hands, so the test does too.

Comparison

Each block against its neighbour; error = difference ÷ √2

Two adjacent blocks sit next to each other in time, so the player is the same person in both. For two independent estimates with the same error, the difference between them is √2 times as noisy as either.

Reported

The 95th percentile of that error

The bad case, not the typical one.

2.3 Finding 1: within one sample, the formula holds

The textbook formula assumes every hand is an independent draw, and the obvious objection is that they are not: the same session, the same table, the same opponents, the same mood. If that objection bites, every margin on this site is too narrow and every finding is overstated.

It is testable without any modelling. Take a player’s own hands and deal them alternately into two halves — same player, same weeks, same tables, every other hand. Both halves describe the same stretch of play, so if consecutive hands carried shared information the two would disagree by more than the formula allows.

They do not. Measured over 7,453 players with 20,000 to 400,000 hands each, the gap between halves, divided by what the formula predicts it should be:

StatisticPlayersMeasuredFormula saysInside the 95% interval
PFR18,0001.0021.00095.1%
CBet flop14,0001.0051.00094.8%
Fold vs CBet — flop7,4530.9981.00095.1%
Fold vs CBet — turn6,0720.9961.00095.3%
Fold vs CBet — river1,5090.9961.00094.6%

A ratio of 1.000 and 95.0% would be the formula exactly right. It holds flat from 400 opportunities to 30,000 and more, so there is no correction to apply at any sample size, and none is applied.

2.4 Finding 2: players drift, and that never averages out

What a block of hands is compared against decides what gets measured.

one block of handscompared with the blocknext to it — same eracompared with a lifetimeaverage — spanning yearsmeasures how precisethe estimate ismeasures that, plus howmuch the player changed
Both comparisons are arithmetic on the same hands. Only one of them answers the question the ± column is asking.

The difference is not small.

Score each block against the player’s lifetime average instead of its neighbour and the error stops falling: VPIP went from ±2.76 to ±2.58 points between 10,000 and 20,000 hands, where pure sampling predicts an improvement of √2.

That floor is not noise. These accounts span years, and players change — they move stakes, adjust, get better or worse. What is left is a floor:

VPIP±2.1CBet flop±8.70246810pointswhere the error settles at 20,000 hands, and stays
However many hands you add, this is where each one stops. The two are not the same distance from the truth, and c-bet is four times further.

So there is no fixed value to converge on: the target moves while you are measuring it.

The same split that settles Finding 1 separates the two effects cleanly. Halve a player’s hands by alternating HAND and the ratio is 1.00; halve the same players by alternating DAY and it rises to 1.67 for PFR and 1.28 for c-bet, growing with volume. Nothing is correlated within a sitting. Months are.

2.5 The band — how close is close enough

The question it answers

How near does a figure have to land before it is worth quoting?

A point means more on some statistics than on others. Players spread across 16 points of VPIP but only 5 of 3Bet, so missing by one point on 3Bet is three times the miss.

Every statistic gets its own band, sized to how far apart players actually sit on it.

  1. 1

    Measure the spread.

    How far apart the middle half of the field is on that statistic.

  2. 2

    Split it evenly.

    Size the band so about one player in ten lands inside it.

  3. 3

    Keep it reachable.

    If a rare spot would need more hands than anyone plays, widen the band, up to about a quarter of the field.

Thirteen statistics fit step 2 as they are, thirteen needed widening and twelve hit the limit. The table shows which.

The rule could be afforded — 13 of 38

IQR ÷ 10.9 — about a tenth of the field inside the band, and a sample that reaches it exists.

StatisticField spreadBandInside itChances
AF0.80±0.1013%7k
3Bet4.83±0.5011%9k
PFR7.38±0.7511%3k
AFq9.29±1.0012%5k
RFI10.67±1.0010%5k
Limp11.41±1.2512%197
Steal15.40±1.5011%5k
VPIP16.36±1.5010%1k
Fold vs Steal16.75±1.7511%9k
Cold Call18.36±1.7510%4k
Limp-Call20.03±2.0011%10k
Limp-Fold20.09±2.0011%9k
Call vs Steal22.11±2.2511%5k

Widened until a real sample could reach it — 13 of 38

The rule would have demanded hundreds of thousands of hands, so the band opens out to what 15,000 buys.

StatisticField spreadBandInside itChances
Limp-Raise3.86±0.5014%16k
Check-raise F5.21±0.7516%15k
3Bet vs Steal5.95±0.7514%14k
WTSD5.48±0.8016%16k
WWSF6.20±0.8014%18k
WWST5.16±1.0021%16k
Raise vs CBet5.90±1.0018%15k
Iso9.63±1.2514%13k
W$SD8.40±1.5019%15k
Fold vs CBet10.41±1.7518%15k
CBet flop14.44±1.7513%14k
Fold vs 3Bet17.51±2.7517%14k
W$SD Bet River14.54±3.0022%19k

Stopped at the cap, and still expensive — 12 of 38

Widening halted at 0.231 × IQR. Past that a band covers a quarter of the field and stops separating anybody.

StatisticField spreadBandInside itChances
4Bet Range1.62±0.3020%36k
Squeeze3.44±0.7524%23k
4Bet5.08±1.0021%41k
WWSR5.21±1.0021%23k
Probe Bet6.01±1.2522%204k
Fold vs CBet T8.31±1.7523%74k
Fold vs CBet R10.05±2.2524%195k
CBet turn11.35±2.5024%24k
CBet river11.52±2.5023%95k
Bet-Fold Flop12.71±2.7523%44k
Fold vs Squeeze12.77±2.7523%785k
Fold vs 4Bet14.59±3.0022%72k

Spread is the middle half of the 6-max field, noise removed. “Inside it” is the share of that field the band covers. “Chances” is what it takes to reach the three-quarter notch on the confidence bar, counted in chances at the statistic rather than hands, since how many hands that is depends on how often the spot comes up. What a value is worth prints the band beside the hand count for the fifteen headline statistics.

2.6 What the ± column is

StatPlayerFieldvolume-matchedPlayer vs field±n
VPIP29.424.4+5.0±1.712,000
4Bet6.25.0+1.2±1.6900

Two rows of the radar readout, with the winners column dropped. Illustrative figures — but each ± is the formula below run on the n beside it. The second gap is smaller than its own margin, so it is shown in amber.

How to read it

Every stat on a player page shows a gap: the player’s number minus the field’s number. The ± next to it answers one question:

If this player were actually just an average member of the field, how far from the field’s number could their observed rate land — on luck alone — 95% of the time?

If the gap they actually have is bigger than that, luck doesn’t explain it.

The rule

If the gap is smaller than the ±, the gap is not evidence.

How it is computed

Each number carries its own margin:

margin = 1.96 × sqrt( p (1 - p) / n ) ─┬── ────────┬──────── 95% level textbook error

n

the player’s opportunities for this statistic — not their hand count

A 40,000-hand player might have 26,000 fold-vs-3-bet spots and 900 4-bet spots; those two rows deserve very different margins and get them.

p

the field’s rate for this statistic — not the player’s

The question is “could someone who really matches the field have drawn this?” Using the player’s own rate shrinks the interval exactly where they are most extreme, which is where it is needed most.

So the only thing making which statistic you are looking at change the answer is p (1 - p), which is largest at a coin flip and falls away towards either end:

Statisticn = 200n = 500n = 2,500n = 10,000
VPIP±6.0±3.8±1.7±0.8
PFR±5.3±3.3±1.5±0.7
Fold vs CBet±6.9±4.4±2.0±1.0

The order never changes. Fold-vs-c-bet sits nearest a coin flip, so it always needs the most evidence; PFR is furthest from one of the three and always needs the least. Every row improves at the same rate, because every row is the same curve at a different height.

2.7 Worked examples

Real numbers from a live profile, arithmetic shown. Each one runs the formula above and holds the answer against the gap.

example 1n = 606,655+5.0REALexample 2n = 2,000+1.6AMBERexample 3n = 50+20.0REAL05101520percentage points
One scale for all three. The margin is the band; the gap is the marker. Outside the band the number prints; inside it, it greys.

Example 3 has one fortieth of example 2’s sample and clears its margin, where example 2 does not clear its own. Neither the gap nor the sample size decides it alone; only the two against each other.

1. A high-volume player’s VPIP — comfortably real

VPIP player 29.4% field 24.4% gap +5.0 points n = 606,655 textbook error sqrt(0.244 × 0.756 / 606,655) = 0.000551 the same, in points × 100 = 0.055 95% margin 1.96 × 0.055 = ±0.11 points gap +5.0 against margin ±0.11 — the gap is bigger -> REAL

2. The same statistic on 2,000 hands — not evidence

VPIP player 26.0% field 24.4% gap +1.6 points n = 2,000 textbook error sqrt(0.244 × 0.756 / 2,000) = 0.00960 the same, in points × 100 = 0.96 95% margin 1.96 × 0.96 = ±1.88 points gap +1.6 against margin ±1.88 — the gap is smaller -> AMBER

3. A big gap on a tiny sample — real

CBet flop player 80.3% field 60.3% gap +20.0 points n = 50 textbook error sqrt(0.603 × 0.397 / 50) = 0.0692 the same, in points × 100 = 6.92 95% margin 1.96 × 6.92 = ±13.56 points gap +20.0 against margin ±13.56 — the gap is bigger -> REAL

2.8 What this does not cover

  • The field’s rate is measured on the pool you filtered to. Change the room or the stake and p changes with it, so the same gap on the same player can clear its margin in one pool and not in another.
  • It is about precision, not truth. On enough opportunities the margin falls below a tenth of a point, and a gap clearing it means these hands are unlikely to have produced it by chance. It does not mean the player still plays that way; that is Finding 2.
  • The verdict panel uses a different, stricter gate. Being flagged unusual takes both a distance from the player pool and two standard errors, on at least 30 observations.

Drift measured on 40 PokerStars accounts (~4 million hands) and re-run 2026-09-27 on 200 accounts of 400,000+ hands each. Independence within a sample measured 2026-09-28 on 7,453 accounts of 20,000 to 400,000 hands, split by alternating hand. Pool, field and winner baselines are rebuilt nightly from the full fact table. See also the statistics guide for what each individual number means.