Model comparison

This page is a live experiment. Rather than pick one algorithm and present it as the truth, several run side by side and are scored against what actually happened.

Every model here is fitted on earlier seasons only and then scored on matches it has never seen — walk-forward validation, season by season. They are judged on how well their probabilities are calibrated rather than on how often they name the winner, because a model that is right for the wrong reasons is not much use.

Bookmaker prices sit on the same board, on the same matches, as a benchmark none of these models is entitled to beat. The aim is not to crown a winner. It is to find out where each set of assumptions holds and where it breaks.

This season so far

2026/27 — 36 matches scored, every model against what actually happened, using parameters fitted on prior seasons only. 42 have been played. A match counts here only where every model priced it, so the figures can be compared at all. Two things leave a played match out: the models that read rolling form have no opinion on a season's opening fixtures, and a match the corpus carries no market prices for is one the bookmaker benchmark cannot price.

ModelLog loss 95% intervalMatches
Poisson 0.9563 36
Dixon-Coles 0.9566 36
Elo 0.9774 36
Adaptive Poisson (Model 6) 0.9780 36
Ensemble (Model 7) 0.9799 36
Bookmaker 0.9803 36
Gradient boosting (Model 5) 1.0090 36
Logistic (Model 4) 1.0351 36
Baseline (Model 0) 1.0861 36
League position (Benchmark 2) 1.1511 36
Always home undefined — 36

The bars are 95% intervals on the log loss, drawn against one shared scale. Where they overlap, this season has not yet separated those models — a season is a small sample, and the ordering above should not be read as a ranking until the bars pull apart. The comparison over 5,258 historical matches is below, and is the one to trust. A model that assigns zero to something that happened has no log loss and no interval, which is why always-home shows neither.

Biggest surprises

Results the goal model gave least chance to. Unlike the table above, this needs no sample size.

31 Jul Dundee United 1–1 Rangers we gave it 20.4% why?
09 Aug Rangers 1–2 Hibernian we gave it 11.0% why?
02 Sep Kilmarnock 0–1 St Mirren we gave it 24.2% why?
15 Sep Hibernian 0–1 Kilmarnock we gave it 24.4% why?
20 Sep Celtic 0–1 Rangers we gave it 24.8% why?

All seasons

Scored walk-forward: every model predicts a season using only prior seasons. All rows share one match set, so the figures are comparable.

ModelLog loss BrierRPS Accuracy ECE HECE DECE A Matches
Bookmaker 0.9553 0.5666 0.1944 0.536 0.028 0.011 0.013 5,258
Ensemble (Model 7) 0.9660 0.5737 0.1978 0.533 0.019 0.006 0.019 5,258
Adaptive Poisson (Model 6) 0.9685 0.5754 0.1987 0.528 0.014 0.003 0.017 5,258
Elo 0.9712 0.5768 0.1993 0.529 0.028 0.010 0.033 5,258
Poisson 0.9765 0.5809 0.2013 0.519 0.018 0.004 0.019 5,258
Dixon-Coles 0.9766 0.5809 0.2013 0.519 0.016 0.003 0.020 5,258
Logistic (Model 4) 0.9835 0.5846 0.2019 0.531 0.025 0.008 0.021 5,258
Gradient boosting (Model 5) 0.9955 0.5928 0.2058 0.520 0.041 0.011 0.028 5,258
League position (Benchmark 2) 1.0079 0.6021 0.2112 0.515 0.023 0.007 0.024 5,258
Baseline (Model 0) 1.0687 0.6468 0.2329 0.437 0.005 0.005 0.001 5,258
Always home undefined 1.1263 0.4446 0.437 0.563 0.237 0.326 5,258

A log loss of undefined is not missing data. A model that assigns zero probability to an outcome that then happens has infinite log loss, and clipping it would report the clipping constant rather than the model. Judge models by log loss, Brier, RPS and calibration — never by accuracy alone: the baseline and always-home score identically on accuracy because both always pick a home win, yet they are far apart on Brier.

Log loss by season

Lower is better. Each point is one season, predicted from prior seasons only. Dixon-Coles sits underneath Poisson — the two differ by at most 0.0016 in any season, so its line is hidden rather than missing. Click a legend entry to hide that model and see what is beneath it.