Methodology

How the numbers on this site are produced, what they mean, and what they cannot tell you. Organised by the question rather than by the machinery.

Where do the numbers come from?

Every match result since the 2000/01 season — the year this league went from ten clubs to twelve, and the year match statistics begin — from football-data.co.uk. That is 5,891 played matches across 27 seasons. Each file is downloaded, checksummed and archived before anything reads it, so a figure on this page can always be traced back to the bytes it came from.

Remaining fixtures come from the official SPFL schedule rather than being derived from the league format. The format is exactly right about which clubs meet and wrong about who is at home in 31 of 186 fixtures, and home advantage is a substantial term in every model here.

What do the models do?

Poisson
Estimates how many goals each club scores and concedes against an average opponent, then treats a match as two independent goal counts. Every scoreline gets a probability; the three outcomes are the sums.
Dixon-Coles
Poisson with a correction for low scores. On this league it changes almost nothing, and we publish it anyway because the null result is the finding.
Elo
One rating per club, moved after every match by the result and the margin. Simple, and it updates as a season unfolds.
Adaptive Poisson
Poisson that watches the season in progress. Club strengths start where prior seasons put them and are pulled toward what this season shows, slowly — the season only outweighs the prior around match 24 of 38. It is the strongest model here that does not use bookmaker prices.
Logistic regression and gradient boosting
Learned from 23 rolling form features rather than from goals directly. They carry real signal and still lose to the football models, which is worth knowing.
Baseline
Predicts the historical home/draw/away frequencies for every match, knowing nothing about the clubs. It is the floor every other model must clear.

How good are they?

Measured, published, and not flattering. Every model is scored on the same matches by log loss, Brier score, ranked probability score and calibration error — never by how often it picks the winner, because a model can be right often and badly wrong about how sure it was.

Bookmaker prices beat every model we have, and that comparison stays on the site because hiding it would make the rest less believable. See the full comparison →

What does the 90% range mean?

The season is simulated ten thousand times. The 90% range is where nine in ten of those simulated seasons put a club's final points, and the measured coverage is 88% — that is, across 24 completed seasons, the real total landed inside the range that often. We publish the measured figure rather than the nominal one.

Each club is given a season-long strength offset, drawn once per simulated season and held across all its remaining fixtures, because a club that is better than its rating is better for the whole campaign rather than match by match. The size of that offset was measured from goal residuals over 293 club-seasons, not chosen to make the coverage look right.

What this cannot do