Methodology
How the numbers on this site are produced, what they mean, and what they cannot tell you. Organised by the question rather than by the machinery.
Where do the numbers come from?
Every match result since the 2000/01 season — the year this league went from ten clubs to twelve, and the year match statistics begin — from football-data.co.uk. That is 5,921 played matches across 27 seasons. Each file is downloaded, checksummed and archived before anything reads it, so a figure on this page can always be traced back to the bytes it came from.
Remaining fixtures come from the official SPFL schedule rather than being worked out from the shape of the league. The format settles which clubs meet and how often, but not who is at home: each pair meets three times before the split, so one club hosts twice and only the schedule says which. Set side by side in August 2026, deriving them put the wrong club at home in 31 of the 186 fixtures then remaining. Home advantage is a substantial term in every model here, so hosts are looked up and never inferred.
What do the models do?
- Poisson
- Estimates how many goals each club scores and concedes against an average opponent, then treats a match as two independent goal counts. Every scoreline gets a probability; the three outcomes are the sums.
- Dixon-Coles
- Poisson with a correction for low scores. On this league it changes almost nothing, and we publish it anyway because the null result is the finding.
- Elo
- One rating per club, moved after every match by the result and the margin. Simple, and it updates as a season unfolds.
- Adaptive Poisson
- Poisson that watches the season in progress. Club strengths start where prior seasons put them and are pulled toward what this season shows, slowly — the season only outweighs the prior around match 24 of 38. With no matches played it is Poisson exactly, not approximately.
- Ensemble
- A weighted blend of Adaptive Poisson, Elo and the logistic model, with the weights themselves learned from earlier seasons. It exists because those three are wrong about different matches: early in a season they genuinely disagree, and the blend is worth most exactly there. Bookmaker prices are deliberately left out — an ensemble given the market becomes the market, which we measured rather than argued about.
- Logistic regression and gradient boosting
- Learned from 23 rolling form features rather than from goals directly. They carry real signal and still lose to the football models, which is worth knowing.
- League position
- Predicts from where two clubs sit in the table and nothing else — no goals, no ratings, no form. Home advantage comes out at about two places in the table, which is a figure you can weigh against your own sense of the game rather than take on trust. A model earning its complexity has to beat this.
- Baseline
- Predicts the historical home/draw/away frequencies for every match, knowing nothing about the clubs. It is the floor every other model must clear.
How good are they?
Measured, published, and not flattering. Every model is scored on the same matches by log loss, Brier score, ranked probability score and calibration error — never by how often it picks the winner, because a model can be right often and badly wrong about how sure it was.
Bookmaker prices beat every model we have, and that comparison stays on the site because hiding it would make the rest less believable. See the full comparison →
What does the 90% range mean?
The season is simulated 100,000 times. The 90% range is where nine in ten of those simulated seasons put a club's final points, and the measured coverage is 89% — that is, across 24 completed seasons, the real total landed inside the range that often. We publish the measured figure rather than the nominal one.
Each club is given a season-long strength offset, drawn once per simulated season and held across all its remaining fixtures, because a club that is better than its rating is better for the whole campaign rather than match by match. The size of that offset was measured from goal residuals over 293 club-seasons, not chosen to make the coverage look right.
What this cannot do
- No expected goals. There is no xG source for this league we would be willing to stand behind, so none is shown. Shots and shots on target are held for every match and a shot-quality measure is the likely route — but it will not be labelled xG, because it would be a different quantity wearing a trusted name.
- No injuries, suspensions or team news. The bookmaker has these and the models do not, which is part of why the market still leads.
- Cup and European fixtures are invisible. Five of the six league fixtures on 22 August 2026 were postponed, each one involving a club playing a European tie that week; the single pairing with no European commitment went ahead. When postponed matches are rearranged, usually midweek, no model here can see the congestion that follows — each is treated exactly as any other fixture.
- Club strengths are fitted on prior seasons. A summer's transfers go unseen until results accumulate. Adaptive Poisson is what closes that gap as a season progresses, and it does so slowly by design.
- The split compresses predictability. After 33 matches the league divides and only closely matched clubs meet, and every model here gets measurably worse.