What the World Cup Did to the Robot
series :: World Cup 2026 · part 2 of 2
- I Built a Robot Instead of Having Opinions
- What the World Cup Did to the Robot « you are here
The Reveal
Spain won the World Cup. 1-0 against Argentina, which happens to be the exact scoreline the robot had settled on the evening before the final: Spain 1-0, an 18% shot at exact and 48/29/23 on the 90-minute outcome. The final is also the one knockout round where the joker requires no thought whatsoever, there’s only a single fixture to put it on. Priced by Part 1’s expected-points formula, that pick was worth points, doubled to 1.68 by the joker. The robot, in other words, ended its tournament with a 6-point play on the table.
I collected 3 of them.
Here the cruel poetic irony that closed Part 11 came to collect its interest. My champion pick had to be locked in by the knockout stages, and at lock time the robot’s favourite was Argentina, so Argentina it was. The robot then spent the following fortnight quietly defecting to Spain, and by the 8th of July it had the pair of them in the final with Spain on top. The scoreline I could still change right up to kickoff, so Spain 1-0 went in and banked its 3 points for exact. The joker I could not: it was locked to Argentina, because that’s who I’d predicted, and no amount of the robot’s revised opinion could move it. Three points banked, three forfeited to a lock I’d have undone a fortnight earlier if the league had let me. The robot was right, which is somehow worse.
Two honest sentences before this reads like a highlight reel. On the eve of the final the model gave Spain 64% and Argentina 36% across 100,000 simulations, so even at its most certain it reserved a one-in-three chance of being wrong. A model like this should be judged on the whole season of numbers it produced, not on whether one 64/36 coin landed the right way up.
Part 1 also left an aside dangling: the model was carrying two different finals, the most frequent pairing across all simulations (Spain v England) and the match-by-match modal path (Spain v Argentina). By the 8th of July the two had re-converged on Spain v Argentina, and reality played the path final. Most common destination and most likely route are different questions, but this time they agreed on the answer.
Below is the bracket exactly as the model predicted it on the 2nd of July, before the round of 16 had been played:
And the same bracket after the final, every tie now a result:
The diff between those two pictures is the story of the knockout rounds. The final pairing was called seventeen days early, as were three of the four quarter-final winners. The hole in the script is Brazil-shaped: predicted to reach the semi-finals, removed in the round of 16 by Norway, 1-2. And the champion call was a 54/46 in Argentina’s favour that went the other way. Reality kept nearly the whole script and rewrote the last line.
That’s the result dealt with. The more interesting story is how the odds got there.
The Race
The chart below is the tournament as the model lived it, one run per day. Every kink is real results being re-fitted and re-simulated; the flat stretches are rest days, because a deterministic model asked the same question twice gives the same answer.
Follow Argentina’s line first. Around 14% on the 13th of June, then three group wins (3-0 Algeria, 2-0 Austria, 3-1 Jordan) while the other favourites wobbled, peaking at 27.4% on the 27th of June. That surge is precisely what talked me into the champion lock, and you can watch it fade as knockout reality priced itself back in.
Germany’s line is the cruellest thing the model produced all summer. They won Group E, at one point 7-1 against Curaçao, and hovered around 5% the whole group stage. Then 0.0%, overnight, on the 30th of June: 1-1 with Paraguay in the round of 32, out on penalties. The line doesn’t dip, it stops. Part 1 promised that the 50/50 shootout coin would matter, and here it is mattering.
The Netherlands are the same story on the neighbouring line. Won Group F, drew 1-1 with Morocco, lost the coin. Two group winners, zero knockout wins, both erased by the one number in the entire model that refuses to hold an opinion. Morocco’s reward is the quiet climb to 7.7%, the best outsider odds on the board, which is what knocking out the Netherlands buys you.
France are the opposite shape. A perfect group (3-1 Senegal, 3-0 Iraq, 4-1 Norway) earned them almost nothing, stuck near 7% because their half of the draw was jammed with rivals. Then the bracket opened, and the line climbed to 11.5% on the 30th of June and 17.7% by the 5th of July without France doing anything differently. The draw is a rating too.
And Spain were never the loudest line on the chart, 15.7% at the parameter lock, around 20% in early July, 32% by the 8th, 64% on the eve of the final, 100% at the end. The eventual champions led the race outright for barely a week of it.
So the race shows what each surprise did to the odds. The next chart asks how surprising each one actually was.
One bar per scored group fixture, and the height is the model’s surprise at the outcome, measured in bits:
An outcome the model had at 50% costs exactly 1 bit, one it had at 25% costs 2 bits, and each extra bit doubles the shock. Spain 0-0 Cape Verde was priced at 8%, and bits, the tallest bar of the summer.
Why the logarithm rather than just plotting the probability? Because surprise should add: two 25% outcomes back to back (2 + 2 bits) are exactly as shocking as one 6.25% outcome (4 bits), and the log is the only scale with that property. That additivity is what makes the average bar height one honest number for the whole month, and that number is precisely the log-loss the backtest knobs were tuned on in Part 1, wearing friendlier units.
The average over the group stage was 1.24 bits per match. A clairvoyant scores 0 and a know-nothing shrugging 1/3-1/3-1/3 at every fixture scores on every single bar, so that gap of a third of a bit is the model’s entire edge over guessing, and it is also why the humans in the league stayed within touching distance.
The labelled spikes hold a lovely counter to the instinct that upset means giant-killing: four of the five biggest shocks were draws the model had priced under 20%. The eventual world champions sit at the top of their own upset-o-meter for failing to score against the side ranked 36th of 48 on the ratings table below. The only non-draw in the top five is Ivory Coast 1-0 Ecuador, which the pundits barely blinked at.
All of these shocks were being fed straight back into the two numbers each team carries. Time to look at what they did there.
The Hidden Numbers
Throughout the tournament, the predictions were published and the ratings behind them never were. This chart is the machinery’s internal state: every team plotted by its two fitted numbers, attack (right = scores more) against defence (up = concedes less). The hollow circle is where a team stood in the fit of the 14th of June, two matchdays in; the solid dot is the final fit; the trail between them is the tournament changing the model’s mind. Remember from Part 1 that these are log-scale, so equal steps mean equal percentage changes: a +0.10 move on either axis is , roughly +10.5% expected goals in every future fixture, and a trail of that length is a big deal.
The full list: all 48 teams' ratings and drift, 2026-06-14 to 2026-07-20 (click to expand)
Sorted by overall strength (attack + defence). att/dfn are the latest fit (2026-07-20); each drift column is the change since the 2026-06-14 fit - the trail in the figure, as a number. Ratings are log-scale (+0.10 of attack = ~+10.5% expected goals in every fixture) and centred on the average of ALL ~200 national teams, which is why every side here is positive. Regenerates with every re-run.
| # | team | attack | drift | defence | drift |
|---|---|---|---|---|---|
| 1 | Spain | +1.58 | -0.06 | +1.86 | +0.50 |
| 2 | Argentina | +1.55 | +0.13 | +1.34 | -0.34 |
| 3 | Portugal | +1.47 | -0.02 | +1.33 | +0.12 |
| 4 | France | +1.58 | +0.15 | +1.13 | -0.13 |
| 5 | Brazil | +1.44 | -0.04 | +1.25 | -0.00 |
| 6 | England | +1.58 | +0.19 | +1.07 | -0.43 |
| 7 | Colombia | +1.18 | -0.20 | +1.46 | +0.27 |
| 8 | Netherlands | +1.51 | +0.07 | +1.00 | -0.06 |
| 9 | Belgium | +1.44 | +0.08 | +1.03 | -0.05 |
| 10 | Morocco | +1.11 | +0.04 | +1.34 | -0.15 |
| 11 | Germany | +1.59 | +0.08 | +0.86 | -0.16 |
| 12 | Switzerland | +1.25 | +0.04 | +1.03 | +0.05 |
| 13 | Japan | +1.20 | -0.01 | +1.06 | -0.11 |
| 14 | Norway | +1.47 | +0.10 | +0.78 | -0.18 |
| 15 | Mexico | +1.05 | +0.09 | +1.12 | +0.03 |
| 16 | Croatia | +1.22 | +0.02 | +0.90 | -0.16 |
| 17 | Uruguay | +0.98 | -0.05 | +1.12 | -0.23 |
| 18 | Senegal | +1.19 | +0.16 | +0.79 | -0.28 |
| 19 | Ecuador | +0.63 | -0.27 | +1.31 | -0.15 |
| 20 | Austria | +1.14 | +0.03 | +0.77 | -0.23 |
| 21 | Australia | +0.80 | -0.24 | +1.11 | -0.09 |
| 22 | Ivory Coast | +0.80 | -0.00 | +1.05 | +0.05 |
| 23 | Turkey | +1.15 | +0.01 | +0.68 | -0.04 |
| 24 | Sweden | +1.23 | +0.08 | +0.55 | -0.12 |
| 25 | Canada | +0.93 | +0.05 | +0.81 | -0.07 |
| 26 | Egypt | +0.87 | +0.13 | +0.87 | -0.09 |
| 27 | DR Congo | +0.72 | +0.12 | +1.00 | +0.04 |
| 28 | Algeria | +1.08 | -0.04 | +0.64 | -0.24 |
| 29 | Iran | +0.81 | -0.21 | +0.90 | -0.03 |
| 30 | United States | +1.07 | +0.00 | +0.64 | -0.10 |
| 31 | Paraguay | +0.65 | -0.19 | +1.06 | +0.25 |
| 32 | Scotland | +0.82 | -0.12 | +0.83 | -0.07 |
| 33 | South Korea | +0.81 | -0.21 | +0.81 | -0.04 |
| 34 | Czech Republic | +0.92 | -0.07 | +0.57 | -0.10 |
| 35 | Ghana | +0.57 | -0.08 | +0.88 | +0.28 |
| 36 | Cape Verde | +0.63 | +0.13 | +0.72 | +0.19 |
| 37 | South Africa | +0.56 | -0.03 | +0.76 | +0.15 |
| 38 | Bosnia and Herzegovina | +0.67 | +0.02 | +0.50 | -0.11 |
| 39 | Jordan | +0.74 | -0.11 | +0.38 | -0.16 |
| 40 | Panama | +0.55 | -0.15 | +0.56 | +0.10 |
| 41 | Tunisia | +0.66 | -0.07 | +0.45 | -0.47 |
| 42 | Saudi Arabia | +0.40 | -0.15 | +0.62 | -0.07 |
| 43 | Uzbekistan | +0.61 | -0.09 | +0.35 | -0.63 |
| 44 | Haiti | +0.56 | +0.01 | +0.16 | -0.13 |
| 45 | New Zealand | +0.49 | +0.16 | +0.20 | -0.41 |
| 46 | Iraq | +0.36 | -0.17 | +0.32 | -0.38 |
| 47 | Qatar | +0.51 | -0.17 | +0.04 | -0.30 |
| 48 | Curaçao | +0.22 | -0.06 | +0.05 | -0.12 |
The first thing to notice is what didn’t move. Brazil finished the World Cup at -0.04 attack, -0.00 defence from where they started it, despite being knocked out in the round of 16. Six matches cannot outshout a decade of evidence, and WC_BOOST = 3 is exactly how loud the tournament is permitted to shout. This is a feature: a model that re-ranked its elite after every upset would be a mood ring, not a model.
The champions are the exception that proves the weighting. Spain’s defence rose +0.50, the biggest single move at the top of the table, and it was earned the hard way: one goal conceded in eight matches. In model terms, every attack that now faces Spain gets divided by an extra . England drifted the other way, attack +0.19 but defence -0.43, which is what reaching a semi-final while conceding in nearly every match looks like on a log scale. The 6-4 third-place game against France did neither team’s defence rating any favours.
The genuine re-learnings live in the middle of the table, where six matches can change the verdict. Norway rose (a second-place group finish behind France, then Brazil removed in the round of 16), Senegal’s attack climbed +0.16 on the strength of a 5-0 over Iraq, and Cape Verde improved on both axes for holding the champions to that 0-0. Going the other way, Tunisia’s defence collapsed -0.47 after shipping twelve goals in three defeats, and Uzbekistan’s fell further still.
This chart is also the answer to the changelog’s most persistent question: why did the prediction for X move when X didn’t even play? Because there is no X in isolation. Every re-fit re-centres all ~200 teams against each other, so Paraguay holding Germany changes what every past result against Paraguay was worth. The model is one organism, not 48 independent dials.
Which raises the only question that actually matters: were the numbers it published any good?
The Report Card
Part 1 ended by asking whether a robot that promised itself 60 points in the group stage would actually collect them. Here’s the homework, graded.
The amber curve is simple: the points genuinely banked, 3 for an exact scoreline, 1 for a correct outcome, accumulating match by match. The grey curve is the unusual one, and it’s the model’s promise to itself. Before each fixture the model knows both its pick and its own probabilities, so it can price what the pick is worth using the expected-points formula from Part 1:
A pick with a 12% chance of being exact and a 55% chance of the right result is worth points, on average, before a ball is kicked. Sum those per-fixture worths across the group stage and you get grey.
The verdict, over 70 scored fixtures: 44 correct outcomes, of which 8 were exact scores, so points banked, against 60 promised. It said its picks would come off 58% of the time; they came off 63%.
It’s worth being precise about what that does and doesn’t prove, because it separates two ideas that are easy to conflate. Accuracy is how many you got right, and it is substantially luck. Calibration is whether reality matched your stated confidence, and it’s the thing a model actually owes you. If the probabilities had been bravado, amber would have sagged below grey; had they been falsely modest, it would have ridden above. Promising yourself sixty points and collecting exactly sixty is the real flex; clairvoyance was never on offer.
And the 8 exact scores from 70 (11%) is not the failure it might look like. It’s the humility the scoreline grid promised back in Part 1: the best available exact pick lands about one time in six, and it did.
Provenance, for the suspicious (click to expand)
Every prediction on this chart was frozen the moment its fixture was played, read back from timestamped snapshots written after each model run, so there is no retro-fitting: the grey curve could not be quietly improved after the fact. Two of the 72 group fixtures were played before the very first snapshot existed and are excluded rather than reconstructed. Knockout fixtures aren’t included because the snapshots only stored per-fixture predictions for the group stage, an oversight I’d fix next time.
Calibration is one half of how the model updates. The other half is blunter: reality simply overwrites the dice.
Reality Overwrites the Dice
Once a match is played it stops being simulated at all. Every one of the 100,000 tournaments thereafter contains Mexico 2-0 South Africa as historical fact; those points can never be un-won. Formally, every number the model published was always , and the conditioning bar is doing more of the work with every matchday: by the semi-finals a “simulation” is mostly transcript with a little dice at the end, which is why group probabilities march to 100% and champion odds squeeze towards 0 or 100. The twelve panels below are that argument, drawn twelve times: the model’s confidence in its group-winner pick (solid) and runner-up pick (faint) after each day’s run.
Learn one cell and you’ve learned all twelve. Flat segments are rest days, every step is a matchday being fitted, and every line is one-way, because information only accumulates; confidence never retreats. The hollow circles are the most honest moments a forecaster has: not “more sure about the same pick” but “it’s not them any more”. Group D carries the cleanest example, the pick flipping to the United States on the 13th of June after their 4-1 over Paraguay, and from there it locked first of all twelve groups. It was also, for balance, the only group where my six scoreline picks banked exactly nothing: the model knew who was going through before anyone else and still couldn’t price a single match there. Certainty about the destination, useless on the route. Group K is the other extreme, a mid-stage flip and no lock until the final matchday, with Colombia only decided by the last round of fixtures. And in nearly every cell the faint line locks later than the solid one, because second place is genuinely the harder call.
The very last run of the season, after the final, was a simulation with no dice left in it: 100,000 identical tournaments, one of them real.
What It Never Knew
A final list, stated with some pride. The robot knew no injuries, no suspensions, no team news, no tactics, no weather, and no narratives. It read a decade of scorelines and not one news article, and it still won the league: 120 points, twenty-one clear of second place, with 72 of the tournament’s 104 fixtures called right in some form and 16 of them on the exact scoreline. It also placed 225th on the platform’s global leaderboard - 197th if you rank by points alone, since a crowd of us tied on 120, and after a summer of shootout coins I know better than to respect a tiebreak.
I should be honest about my own contribution to all this: I typed in whatever the robot said and got on with my week. Part 1 admitted that football offers me little entertainment value personally, and a full World Cup has done nothing to change that. It turns out you can win a football leaderboard without acquiring the slightest interest in football, and I’d be lying if I said that wasn’t my favourite result in this whole post.
Things I’d try next time, one sentence each. The Dixon-Coles low-score correction2, because four of the five biggest shocks being cheap draws is exactly the bias it fixes. Strength-weighted penalty shootouts, although “I tested it and nothing changed” remains the pre-registered reply, even after watching the coin delete Germany and the Netherlands in a single evening. Fitting the ratings on shot-based xG rather than goals, which is the credible upgrade but needs data that simply doesn’t exist for most international sides. And per-venue altitude and travel effects, for a tournament spread across three countries.
So what did I actually add? The honest answer is: not the maths. The model is proudly 1982 (Maher’s attack/defence Poisson), the recency weighting is 1997 (Dixon & Coles), and the simulate-it-100,000-times architecture is exactly what the academic consortium of Groll, Zeileis and colleagues runs, down to the same simulation count3. Their pre-tournament forecast had Spain on top at 14.5%, and the favourite duly won; the robot’s list on the opening weekend also had Spain first, at 15.7%, before it spent a fortnight infatuated with Argentina. On the one scoreable comparison of the summer, the professionals and the hobbyist agreed more than either would like to admit.
What’s genuinely mine is the two layers around the maths. The decision layer, because no academic has to hand one scoreline to a league admin by Friday: picks priced under a 3/1 scoring rule, a joker spent by expected points, draws avoided not by opinion but by geometry. And the ops layer: a fixed seed, a timestamped snapshot after every run, a changelog line attributing every moved number to data or a knob, parameters locked on the 13th of June, and pre-match predictions frozen so the report card couldn’t cheat. Neither layer is a better model. Both are what let a 1982 model be accountable in 2026.
That, I think, is the meta-lesson of the whole exercise: the engineering mattered more than the maths. Calibration, not clairvoyance, is what a good model owes you, and it’s a debt you can only prove you’ve paid if you wrote everything down at the time.
The robot is switched off now. It retires with a league title, exactly the group-stage points it promised itself, one perfect final it was never allowed to double, and no opinions whatsoever about Euro 2028.
On that, for once, we agree.
Footnotes
-
Dixon & Coles (1997), “Modelling Association Football Scores and Inefficiencies in the Football Betting Market”, JRSS C 46(2) ↩
-
https://www.r-bloggers.com/2026/06/football-meets-machine-learning-forecasting-the-2026-fifa-world-cup/ ↩