same match, different number

Why two models disagree on the same match

Feed two honest models the same fixture and they hand back different numbers. The disagreement is not a bug, it is where the judgement lives. A tour of the choices that move the answer.

21 July 2026 · 6 min read · James Frewin

Cover illustration for Why two models disagree on the same match

Feed the same fixture to two models built by two careful people and you will get two different numbers back. Not wildly different, usually, but different enough to argue over: 48% here, 53% there, on the same match, from broadly the same data. The instinct is to assume one of them has a bug. Almost always, neither does. The gap is not the models failing to agree. It is where the human judgement lives, out in the open, in numbers you can actually poke.

The rates are the whole ballgame

Every scoreline model starts in the same place: a scoring rate for each side, the expected goals it thinks each team will create over ninety minutes. Almost everything else is downstream of those two numbers, which is why they are by far the biggest lever. Nudge one team from 1.4 expected goals to 1.7 and the entire distribution of results shifts with it.

And that rate is not read off a chart. It is a judgement. How much does a key injury cost? Does this side travel well? How do these two ways of playing collide, does one team’s press feed the other’s counter? Two modellers can study the identical form guide and honestly land a few tenths of a goal apart, and a few tenths is enough to move a win probability by several points. This is the single largest source of disagreement, and it is not going away, because it is the part that is actually about football.

How much to trust a short run

The second dial is priors and shrinkage: how much a model believes a small, recent sample. A team wins four on the bounce. Is that who they now are, or a hot streak sitting on top of ordinary quality? A model with strong shrinkage pulls that run back toward the long-run baseline and stays sceptical. A model with light shrinkage takes recent form closer to face value. Neither is wrong in principle. They are different bets about how much signal a handful of matches carries, and they will price the same in-form team differently every single time.

Two honest models disagreeing is not a contradiction to resolve. It is two defensible readings of the same uncertain thing.

The quieter dials

Three smaller choices move the number in ways people rarely notice.

Home advantage. Everyone builds in an edge for the home side, but not the same edge. A model calibrated on packed league grounds and one tuned for neutral tournament venues will disagree before a ball is kicked, purely on this.

How the two teams’ goals are linked.Goals are not fully independent: a game that opens up tends to open up for both sides at once. A model that builds in that correlation puts a little more weight on the draw and on low-scoring games than one that treats each team’s goals as separate. That one choice quietly reshapes the draw probability, the hardest outcome in football to call.

Recency weighting. How fast does an old match stop mattering? A model that forgets quickly reacts hard to the last month; one with a longer memory rides through a rough patch. Same history, different half-life, different number.

A five-point gap is the system working

Put all of that together and a difference of around five percentage points between two good models is not a red flag, it is the normal texture of honest forecasting. If anything, two models agreeing to the decimal is the thing to distrust, because it usually means they are sharing assumptions rather than reasoning apart.

The betting market is worth a glance here, as one more number in the mix, a rough consensus with money behind it. But it is a forecast too, with its own assumptions (and a margin baked in), not a referee that tells the others they are wrong. Treat it as another opinion, not the answer.

The illustration below makes the point with a made-up match: two models, the same fixture, splits that differ by a handful of points across home, draw and away. Both are perfectly reasonable. That is what disagreement looks like when everyone is being honest.

Model A46%27%27%Model B51%24%25%Home winDrawAway win
Illustrative only: one made-up fixture, two honest models. The splits differ by a few points across home, draw and away. Both are perfectly defensible.

So there is no single true number waiting to be uncovered, only a set of defensible ones, each carrying its author’s judgement about rates, samples, venues and time. A model you can trust is not the one that claims to have settled the question. It is the one that shows you the dials it turned, so you can turn them yourself and see where your own honest number lands.