How the model actually did.
Every match joined to the model's morning-of prediction, then measured against the result. The structure was excellent, a match for the sharpest bookmakers and ahead of them in the knockouts: who wins, the favourite calls, the Golden Boot. There is one honest weakness, and it's goals. Here is all of it, nothing hidden.
Results vs the pre-tournament board
BankedFavourite accuracy climbs as the tournament sharpens. The model reads a two-good-teams knockout far better than a noisy group game. Goal markets (right two columns) tell the other story: model under-calls in every round.
In a 48-team field even the favourite is under 50% to reach the semis, so the honest read is the ranking, not any per-team threshold. The board, ordered by probability of winning, against where each side finished:
The top four by probability were the exact four semi-finalists. Spain (#1) won it, Argentina (#2) runner-up; only France and England swapped 3rd and 4th. The board also rated Brazil below their reputation, at #5 rather than the popular #1, and Brazil went out in the Round of 16.
Biggest over-performers the model rated low: Norway (reached QF, P(win) 1.3%), Switzerland (QF, 1.6%), Morocco (QF, 2.4%). Tournament football keeps its surprises. The board priced them as long shots, correctly, and a few of the long shots came in.
Golden Boot
BankedThe model's top two pre-tournament picks were the two top scorers, and every one of its top six board picks scored. Kane sat top of the pre-tournament board and came back toward the pack as the tournament played out.
Board shows the pre-tournament probability of finishing top scorer; Mbappé and Messi sat just off the top of it and both climbed as their goals arrived. Every one of the board's top six picks scored at the tournament (Haaland 7, Kane 6, Lautaro, Havertz and Undav three each, Torres one). The board leans on proven scoring volume, a deliberately cautious starting point.
Over 2.5 & BTTS, the real miss
MissThis is the finding that matters, and we are not hiding it. The model under-called goals by around 10 points on Over 2.5 and around 12 on Both Teams To Score, in the group stage as well as the knockouts. Part of that is a genuinely high-scoring tournament (see below); part of it is ours to fix.
Two things to separate. One, level: our goals baseline is tuned to a typical World Cup, so a hot one runs past it. Two, shape: we under-estimated how often both teams score, independent of the total. The level is partly the tournament, as the history below shows. The shape is ours, and it is the focused piece of work we are prioritising next.
Was it really that hot? Yes. This was the highest-scoring men's World Cup in at least twenty years, above every edition since 2006. Our goals baseline sits close to the historical average, so a tournament this open was always going to run past it.
Prior-edition average 2.51, previous high 2.69 (2022). 2026 came in at 2.96, a full quarter-goal above the previous record. A model anchored to history was always going to under-shoot a tournament this open on the level, but that does not excuse the shape: both-teams-to-score was the bigger miss, and correlation between the two teams' goals is ours to model better.
Scorelines
MixedExact-score is inherently low-hit, around 11 to 12% is par, and the model landed there. But the low-scoring bias from the goals miss shows up clearly in what it kept picking.
The model reached for 1-0 forty-two times; reality's most common score was 1-1. That one contrast is the goals miss made visible. Scores, both-teams-to-score and draws all move together, so the same focused improvement lifts all three.
Model vs the market
Split decisionThe honest scorecard. On the knockout games where we logged a closing line, the model edged the market on outcomes and trailed it on goals. Across the wider group stage the two were line-ball. Nothing here is cherry-picked.
The honest read. On outcomes the model matched or beat the closing line. On goals it trailed, because even the market under-priced how open this tournament was. Widen the lens to all 63 group games we could price against a market line and the two were line-ball: model and market within a hundredth of each other on log-loss, agreeing on the favourite 98% of the time. Competitive with the market on results, with clear room to improve on goals.
What we change next
See how it looked
The interactive projected bracket the model published through the tournament is preserved. Open it to walk the groups and the knockout tree as it stood.