Machine Learning: The Model

What the ML rating is, and how to use it

You have spent eight chapters learning to read a race the way a person reads it. This chapter is about a second opinion — one that has read more races than you ever will.

Theory: the model

MWP's Expected Rating is generated by a machine-learning model trained on millions of race starts, combining hundreds of data points per horse — recent form, surface and going aptitude, trainer and jockey statistics, gate, class, distance, weight, days since last start, field-relative rankings, and more.

Think of it as a very well-read second opinion.

Where it came from — the Benter story

In 1986 two researchers, Bolton and Chapman, set out the idea every serious racing model still rests on: treat a race as a competition, and have the model give each horse a probability — numbers that add up to 100% across the field, rather than a single "this one wins" guess. A few years later Bill Benter built a computer model on that idea in Hong Kong, and over two decades his syndicate won close to a billion dollars.

Three lessons drive every serious model — MWP's included:

  1. Output probabilities, not point predictions. Not "horse 4 finishes 2nd" — a probability per horse, summing to 100%.
  2. Beat the public, or combine with it. Benter found his standalone model was biased relative to the market. His breakthrough was a second model that blended his ratings with the public odds. The combined model crushed either one alone.
  3. The edge is small but compounding. Benter's edge over the public was about 1–2 percentage points — tiny, but enough to fund a billion-dollar operation when paired with disciplined staking. And it is an arms race: the accuracy a model needed to clear the Hong Kong market rose by roughly 40% over twenty years as everyone's numbers sharpened. An edge is never banked; it is defended.

What the model is good at

The model sees patterns across a volume of races no human can hold in their head. It is consistent — it never has a bad day, never anchors to a horse's reputation, never gets bored in race six. For the baseline — what should we expect, all else equal — it is excellent.

The early racing models were relatively simple. Modern ones are far more powerful — they can weigh hundreds of factors at once and find combinations a person would never think to test. The most reliable ones are ensembles: rather than trust a single model, you run several different ones and blend their answers. That works for a simple reason — different models make different mistakes, and averaging them cancels the errors faster than it cancels the real signal. Benter proved the point with his own figures: his model alone and the public's odds alone were each unprofitable; combined, they made money.

What the model can't see

The model doesn't watch races. It doesn't know the trainer moved stables last week, that the stable jockey was suspiciously replaced, that the horse was boxed in and never got a run, or that today's three front-runners will burn each other off. Those are your lenses.

It also depends entirely on its inputs. A debutant with no race history, or a horse with a thin record on today's surface, gives the model little to work with — treat its number as low-confidence there. And every honest model is built to avoid data leakage — letting the model peek at information that wasn't available before the gates opened (final times, future results, the final odds). A model that leaks looks brilliant in backtests and loses money live. The rule the model is built on: at prediction time it can only see what a bettor could have seen five minutes before the off.

From rating to probability

A raw rating isn't yet usable. To bet, you need race-normalised win probabilities that sum to 1 — horse 4 has 28%, horse 7 has 19%, horse 2 has 14%, all adding to 100%. Betting requires expected value, and expected value requires a probability. You cannot compute "is this a good bet?" from a rating, a rank, or a top pick.

This is the single most important habit in this chapter: the top pick is not the same as the best bet. A 45% favourite at 6/5 returns 45% × 2.20 = 0.99 — negative value. A 14% longshot at 10/1 returns 14% × 11.0 = 1.54 — large positive value. Read the edge, not the top of the rank list.

How good is the rating, really?

That is a fair question to ask of anything you are sold — so here is the honest test, run across the full MWP history. Take the horse carrying the highest MWP rating from its last start in each race:

Now the part that has to be said plainly, because it is the difference between a real claim and a sales pitch: this does not mean you profit by backing it blind. When you actually bet, you pay the margin on every ticket — and a flat bet on the top-rated horse loses money at every price (about −10% on turnover, worse at long odds). Beating the crowd's estimate by 6% is not the same as beating the price, which carries a 15–25% takeout on top.

So why does it matter? Because a single number, with zero judgement applied, claws back most of the takeout and lands you ~6% closer to the market than the market's own consensus. That is genuinely hard to do, and it is the raw material of an edge — the floor you start from before you add a single thing you know. The rest of this course is about converting that floor into profit: skipping the races where you have nothing, betting only the overlays, and playing into the deepest, lowest-margin pools you can reach. The rating gets you to the line. Your judgement steps you over it.

How to use it

Treat the Expected Rating as your starting point, then apply what you know on top. Because horses regress from peak figures, the Expected Rating often looks lower than recent best form — that is statistically correct, not a flaw. The value is in the relative differences between horses.

A real one. In the Dubai Turf (Chapter 12), the rating made Soul Rush the clear second-best in the field — yet the market, transfixed by the odds-on superstar above him, let him drift to fifth choice. The model was never going to pick him over Romantic Warrior; what it flagged was that the price on the second horse was wrong. That gap — where the rating ranks a horse versus where the market ranks it — is the entire job. The rating finds it; you decide whether it's a bet.

The handicapper's edge is identifying which horses will run to or above their best — which is often what winning demands. The model sets the baseline. Your judgement decides who beats it.

[QUOTE: "The goal is not to know who will win. It is to know the price at which each horse is worth backing."]

Practical

Quiz

Up next: pricing — turning all of this into a number, and a fair-odds line of your own.