Results

Every call, frozen before kickoff and graded against what actually happened.

loading…

Predictions are reconstructed from the last pre-kickoff snapshot and nothing else — no post-kickoff row is ever read — so this backfill is hindsight-free rather than a look-ahead. Baselines were fixed before any result was scored.

Calibration — when it said X%, how often did it happen?

The marker shows the middle of each band. A well-calibrated model lands its bar on the marker.

How to read this

Log loss is the penalty for the probability given to what actually happened — lower is better, and it punishes confident wrong calls hard. A verdict only reads beats it when the entire 95% confidence interval on the paired difference sits below zero. Paired matters: model and baseline score the same matches, so testing two independent averages would throw away most of the power and call real differences noise.

The null decides the story. Measured on 40,604 archive matches, the scoreline model beats a league-average Poisson by 0.1151 but beats a book-trivial model — the same machinery fitted to that match's own closing 1X2 — by only 0.0217. Four-fifths of the apparent win was beating a null no bookmaker would ever use. Every card here is scored against the hardest null we have, and where a softer one was used before, the number has been replaced rather than kept alongside.

The full-time column is the market de-vigged. It is not our forecast, and beating a base-rate prior with it says nothing about us — it is the bar. The half-time distribution and the exact scoreline come out of eQt and no book in our tape posts them; those are the only lines here that are genuinely ours.