Machine Learning — Use the Baseline Properly

A consistent estimate is useful even when it needs questioning.

Theory: what the model contributes

A model can compare more historical situations than you can hold in your head. It can apply a method consistently across a whole card, including the horses you find uninteresting. That makes it a useful starting point.

MWP's Expected Rating estimates a horse's likely performance from available form and other inputs. It should be read alongside its definition and the evidence behind the horse. A thin record makes the task different from predicting an established runner on familiar conditions.

A model can be consistent and still systematically wrong. Data errors, changing racing populations and weak coverage can all affect it. The useful question is where its estimates hold up, where they do not, and whether a proposed human adjustment improves them.

A rating, a ranking and a probability are different outputs

An expected rating describes a level of performance. A ranking orders the runners. A win probability describes a chance within the whole field.

Those outputs are related, but one is not a substitute for another. A small expected-rating advantage over one opponent is different from the same advantage over fourteen. The uncertainty around each horse also matters.

A probability model can use the ratings, field and other inputs to estimate race-normalised chances. A model can also combine its own assessment with market information. These are separate modelling choices; do not assume a displayed rating already does all of them.

If the product shows only a rating, do not pretend it has supplied an exact betting probability. If it shows a probability, check the model's description and information timestamp.

The top pick is not the bet

An invented favourite with a 45% chance at 2.20 has expected net return of 0.45 × 2.20 − 1 = −1%. Another runner with a 14% chance at 11.00 has expected net return of 54%.

The arithmetic is easy because the probabilities are supplied. In practice, their accuracy is the difficult part. A ranking disagreement with the market cannot replace that work.

Soul Rush's case in Chapter 12 asks whether a well-rated horse could be overlooked in a field dominated by a famous rival. The rating makes it worth investigating. It does not demonstrate an incorrect price just because the form and market rankings differ.

Know what the model already sees

Do not assume the model is blind to trainer changes, draw, layoff patterns or pace-related information. The available MWP feature documentation includes many such fields, but the exact inputs vary with the model and version.

Likewise, the existence of a field in a database is not proof that it is used well. A coarse surface category or an incomplete trip indicator might miss something specific. That is a reason to investigate a model limitation, not a licence to add a bonus every time you recognise a familiar factor.

A useful manual adjustment names its added information. “I watched the race” is not enough. “The previous run was materially restricted, that restriction is absent from the input I am using, and today's route may differ” is a testable explanation.

Peaks and expectations

Expected ratings often sit below a horse's best recent figure because a peak is not the typical outcome. That can be sensible even for a horse in good form.

It is not a guarantee that every lower model estimate is correct. A young horse may be improving, the historical sample may be sparse, or the conditions may have changed. Compare the model's expectation with your evidence rather than choosing whichever number makes your selection look best.

Readiness is a useful example. If the model already accounts for time since the last run and campaign stage, those facts alone are not new human information. A reliable preparation detail outside those inputs might be.

The body-weight and trials ideas discussed earlier remain research candidates unless a particular product release explicitly documents their collection and use.

What honest evaluation looks like

Test a model on later races it did not train on, using only information available at its declared prediction time. Check data processing too: a trainer aggregate or revised historical record can leak later information without the model itself doing anything obviously suspicious.

Compare on the same races with meaningful baselines. A model beating random selection is a much weaker result than improving on a strong market-based estimate.

For probability forecasts, inspect both scoring performance and calibration. Calibration asks whether events assigned similar probabilities occur at roughly those frequencies over a sufficiently informative sample. A model calling many horses 20% should not systematically see only 10% of that group win.

Calibration alone is not everything. A forecast giving every runner an uninformative average can miss important differences. Proper probability scores, such as a stated version of the Brier score, help compare forecasts across the same outcomes. Cash returns add a separate test involving prices and execution.

We do not use a win-rate statistic for the highest historical rating as proof of the Expected Rating model's performance. They are different tests.

Test the human layer too

Save the baseline before making changes. Record the adjusted figure or probability and the reason. Later compare the original and adjusted estimates on the same races, not only on the bets you remember.

Your adjustments may help with one type of situation and hurt elsewhere. A short run of wins is not enough to know. Avoid slicing the data into ever-smaller successful categories after the fact.

One of the best uses of the tracker is discovering that a favourite personal angle adds confidence without improving prediction. That is uncomfortable information worth having.

Practical: model versus you

Quiz: model claims

Up next: convert a race opinion into a full price line, then decide whether anything is worth buying.