Data Literacy: Stats Aren't Stats
Sample size, confidence, and not fooling yourself
You just spent a chapter learning to trust trainer statistics. This chapter teaches you when not to.
Theory: a stat is not a stat
"This trainer wins 33% first-up." Impressive — until you learn it's 1 win from 3 runners. "This jockey is 0% at the track." Damning — until you learn it's 0 from 2.
A number without a sample size behind it is noise wearing a suit. A percentage without a denominator is a marketing slogan.
Sample size — the most important number on the page
The same headline percentage means completely different things at different sample sizes. The practical hierarchy:
- Under 30 starts — essentially noise.
- 30–100 — suggestive only.
- 100–500 — moderate confidence.
- 500+ — a strong signal.
Two trainers: A wins 8 from 20 first-time geldings (40%); B wins 38 from 200 (19%). The eye is drawn to 40% — but that headline is built on the equivalent of 20 coin flips. B's 19% rests on ten times the evidence and is far more likely to be real. A 40% rate over 10 starts is nearly indistinguishable from a 10% rate over 10 starts: both sit inside pure luck.
Confidence — putting a range around the number
MWP shows a confidence measure alongside its statistics — a Wilson-style interval telling you how wide the true rate could realistically be: SOLID, FAIR, WEAK, NOISE. Read the confidence label before you read the number.
The intuition: a 25% rate from 20 starts could really be anywhere from 6% to 44% — useless. The same 25% from 200 starts narrows to roughly 19%–31% — now you can act. From 1,000 starts it tightens to about 22%–28%. The margin of error roughly halves every time you quadruple the sample. And if the gap between two trainers' rates is smaller than the wider of their intervals, treat them as equal.
How the band is built — and how to read it. The interval is a Wilson score interval: the statistically honest range for a rate given how many starts it rests on. You never need the formula; you need the habit of reading the label before the number. MWP collapses the band into four words:
- SOLID — the band is tight; the number is close to the truth. Act on it.
- FAIR — usable, but hold it loosely.
- WEAK — the band is wide; treat the number as a hint, not a fact.
- NOISE — the sample is so small the rate could be almost anything. Ignore it.
The label already folds in the sample size, so it answers the only question that matters before you trust a percentage: how sure can I be this isn't luck? Read the label first, the number second.
Simpson's paradox — when the aggregate lies
A trend in grouped data can reverse inside the subgroups. A yard shows 30% overall — but split it, and it's 40% in maiden races (mostly short-priced favourites) and 15% in handicaps. If today's runner is a handicapper, the relevant number is 15%, not 30%.
A second classic: a horse 4-from-7 on soft and 1-from-15 on good. "Loves the mud" — except every soft start was off a fresh prep and every good start was deep in a long campaign. The variable doing the work is fitness, not going. Always demand the breakdown by class, surface, distance, going and field size.
Regression to the mean
Extreme performances tend to be followed by ones closer to the average. A horse that runs a career-best 105 off a history of 88–92 is more likely to regress toward 92 than repeat 105. A trainer 5-from-5 over a fortnight on a career 14% will revert. Weight career averages, discount the peaks. A horse that runs 88-89-90-89-91 is a more reliable 90 than one that runs 75-102-78-80-95 — same average, but the public will overpay on the back of the 102.
From our own data — the lesson in our numbers. When we measured trainers' first-up records (Chapter 5), the spread was sample-size discipline in miniature: Jim Goldie −49.6% over ~500 starts (trust it), Annike Bye Hansen −41.2% over 121 (usable, but only just past the line), and the Scandinavian yard from the streak case once reading 80% off five starts (pure noise — his real rate is ~20%). Regression bites the extremes too: a trainer slice showing +80% over 180 starts is a real edge, but it will fall back toward +30–40%, never repeat the 80. Read the number, then read the n underneath it.
Streaks — the gambler's fallacy in reverse
In a large enough dataset, streaks are statistically required. A 15% trainer running 1,000 horses a season will throw a 5-race winning run by chance alone. A streak is real only when it is sustained over months not weeks, across a large sample, with an explainable cause (new assistant, new feed, a key jockey arrangement), and it survives a subgroup check. Three of those four boxes unticked — treat it as noise.
Base rates and not fooling yourself
Start from the base rate, then update. Useful priors: favourites win about 33% globally; class droppers on the same surface 18–22%; first-time starters 8–12%; horses returning from 6+ month layoffs about 8%. Most longshots lose. Most layoff horses lose. Your evidence has to overcome a presumption — it doesn't start from a blank slate.
The fortune-teller's trick
Here is a test to run on yourself. Take any fact about a horse and notice how easily it argues both ways:
| The fact | The bullish read | The bearish read |
|---|---|---|
| Runs again after just a week | Trainer's confident — strike while it's fit | Can't be that well — thrown in to drop its mark |
| Beaten favourite last time | Well-meant, ran into trouble, value now | Exposed as overrated — leave it |
| First-time blinkers | A sharpener; expect improvement | The trainer's run out of ideas |
| Won well last time | In form — follow up | Peaked, mark's gone up, bounce coming |
Every row is true in both columns. Which one you see depends almost entirely on whether you already liked the horse. That's the trap: most people pick the horse first and the reason second, then feel like they did the work. The fix is mechanical: decide what would make you change your mind before you read the form, not after.
The hardest stat to read honestly is the one that confirms what you already wanted to believe. Build the habit: when a number supports your bet, check its sample size harder than when it opposes it. Data literacy is the discipline of being harder to fool, including by yourself. That matters more than knowing extra statistics.
Practical
The same trainer can look like a genius or a journeyman depending on which window you read. Read both panels — same handler, two stories.
Quiz
Up next: the gate draw — where the post position helps, hurts, or lies to you.