grading a forecast
What makes a forecast good? Brier scores and calibration
A forecast is not right or wrong, it is well or badly calibrated. The two numbers that actually grade a probability, explained without the jargon, and how Touchline grades itself with them.
21 July 2026 · 6 min read · James Frewin

Ask most people whether a prediction was good and they check one thing: did it come in? It is the wrong test. A forecast is a probability, and a probability is never simply right or wrong. It is well or badly calibrated, and there is a clean way to measure that. Two ideas do all the work.
Calibration: does 70% mean 70%?
The first idea is calibration. A forecaster is well calibrated if the things they call at 70% happen about 70% of the time, the 30% calls come in about 30% of the time, and so on, all the way across. That is the entire test of honesty. It says nothing about any single match; it is a claim about the long run.
You check it by bucketing a big pile of forecasts by the number attached, then asking how often each bucket actually came true. Plot predicted against observed and a perfect forecaster traces the diagonal.
A 60% call that loses is not a bad forecast.It is supposed to lose 40% of the time. Judge it on one result and you are grading noise. This is exactly why a confident pundit who “called it” tells you nothing, and why a weather forecaster who says 60% and stays dry has not failed.
The Brier score: one number for the lot
Calibration is the picture; the Brier score is the number. It is simpler than it sounds: take the probability you gave, subtract what happened (1 if it did, 0 if it did not), square the gap, and average over every call.
The Brier score is strict in a useful way: it rewards being confident and right, and punishes being confident and wrong harder than it punishes hedging. You cannot game it by always sounding sure, and you cannot game it by always sitting on the fence. The only way to score well is to be genuinely, repeatedly calibrated.
Sharpness without calibration is just confident nonsense. Calibration without sharpness is useless honesty. A good forecast needs both.
That second half is sharpness: how far you dare move from 50/50. Two forecasters can both be calibrated, but the one who confidently says 85% when it is really 85%, rather than mumbling 60%, is more useful. The Brier score quietly rewards that too.
How we grade ourselves
This is the rule the receipts run on. Every settled match gets a Brier score on the win probability the model published before kick-off, never edited after. The scores are averaged in the open, next to the calls, with the misses left in. The average will appear on the receipts once enough matches have settled to score.
So next time someone waves a correct call at you, ask the quieter questions. What number did they put on it, in advance? How do their 70% calls do over a season? Those are the questions a forecast can actually be graded on. Everything else is a story told after the whistle.
More from the journal


