How Do football Prediction Models work?

By ProbaPredict Data Desk · Last reviewed

Most football prediction models estimate how many goals each team is likely to score, then turn those expected goals into probabilities for every scoreline. Ours rates each team's attack and defence from recent results, adds home advantage, uses a Poisson distribution with a Dixon-Coles correction for low scores, and is checked against what actually happens.

Step 1: how are team ratings estimated?

Every team gets two numbers: an attack rating (how many goals it tends to score relative to average) and a defence rating (how many it tends to concede). The model finds the ratings that best explain every result in the training data. The training data covers roughly the last two years of matches (including second-tier leagues so promoted teams have a track record); five seasons of history are used for back-testing and calibration. A team that scores lots against strong defences earns a higher attack rating than one that scores lots against weak ones.

Step 2: why weight recent matches more?

Teams change. We apply time decay: the weight halves roughly every year, so a match from last month counts nearly three times as much as one from 18 months ago. That lets ratings follow form, transfers and managerial changes without overreacting to a single result.

Step 3: how is home advantage included?

The model fits a single home-advantage factor that multiplies the home side's expected goals. In today's fit it is 1.232, a 23.2% boost. Neutral-venue matches get none. See how big home advantage is.

Step 4: how do expected goals become probabilities?

Expected home goals = home attack × away defence × home advantage × a league baseline; expected away goals similarly without the home factor. A Poisson distribution then gives each team's chance of 0, 1, 2… goals, and multiplying gives every scoreline from 0-0 to 10-10. Summing the right cells gives home/draw/away, Over/Under 1.5, 2.5 and 3.5 and the most likely scores.

Step 5: what does the Dixon-Coles correction do?

Plain Poisson treats the two teams' goals as independent, which slightly misprices low scores: real football has a few more 0-0s and 1-1s than it predicts. The Dixon-Coles correction adds a parameter, rho, that adjusts the four lowest scorelines. Today's fitted rho is -0.040.

Worked example

Home attack 1.20, away defence 1.05, home factor 1.12, baseline 1.30: expected home goals = 1.30 × 1.20 × 1.05 × 1.12 = 1.83. Away attack 0.95, home defence 0.90, baseline 1.30: expected away goals = 1.30 × 0.95 × 0.90 = 1.11. Summing the scoreline grid gives roughly home 54%, draw 24%, away 22%, and Over 2.5 about 57%. Fair odds are 100 ÷ each percentage.

How does the model handle uncertainty?

Match predictions use the best-estimate ratings. For season simulations we also allow the ratings themselves to be wrong: each of 10,000 simulated seasons draws slightly different ratings, with more spread for teams we have less data on, plus a little random drift through the season. The amount of spread (k) and how strongly newly promoted teams are pulled toward a typical promoted-team rating are both chosen each month by a back-test.

How do we check the model?

Two ways: every live prediction is graded in public on our accuracy page, and we run walk-forward back-tests that predict past matches and seasons using only data available at the time. A well-calibrated model's 30% outcomes should happen about 30% of the time:

Match calibration (back-test): predicted vs observed, home/draw/away
Predicted bandOutcomesAverage predictedActually happened
0–10%9107.2%6.6%
10–20%5,35116.2%15.7%
20–30%23,13025.6%26.3%
30–40%12,19634.3%33.1%
40–50%7,65644.6%44.1%
50–60%4,17054.4%54.9%
60–70%1,88864.2%65.3%
70–80%82774.3%76.7%
80–90%26883.9%84.0%
90–100%2591.7%84.0%

Source: ProbaPredict analysis of 18,807 matches, updated Fri, 2 Oct 2026. Back-test: each match contributes three outcomes (home, draw, away), predicted using only earlier data.

Season calibration (back-test): title, top-places and relegation chances
Predicted bandTeam outcomesAverage predictedActually happened
0–5%6,0230.5%0.4%
5–15%7329.3%10.2%
15–35%63923.9%21.9%
35–65%52348.6%50.9%
65–85%22275.2%73.9%
85–95%11790.5%94.0%
95–100%23498.9%97.9%

Source: ProbaPredict season back-test of 8,490 team outcomes, updated Fri, 2 Oct 2026. Seasons re-simulated after matchdays 5, 10, 20 and 30 using only results known at the time. Overall Brier score 0.045 vs 0.101 for a naive guess.

Common mistakes when reading model output

What can't the model see?

Because it learns only from results, our model doesn't know about injuries, suspensions, rotation, managerial changes until results reflect them, or motivation late in the season. It also treats every competitive match within a league setting the same way. These are deliberate trade-offs: the model is transparent, consistent and fully reproducible, and its errors are measurable.

Those blind spots are one reason margin-free closing odds, which absorb team news and much more, are slightly more accurate in our back-test. We'd rather publish that than claim an edge we can't show.

Frequently asked questions

What is a Poisson model in football?
A model that treats each team's goals as a Poisson random variable with a given average (expected goals). It gives the chance of scoring 0, 1, 2 or more goals.
What is the Dixon-Coles model?
A refinement of the Poisson model, published by Mark Dixon and Stuart Coles in 1997, that corrects the probabilities of 0-0, 1-0, 0-1 and 1-1 and weights recent matches more heavily.
Does your model use player data or injuries?
No. It uses match results only. That keeps it transparent and consistent, but it cannot react to team news.
How do you know the model works?
We grade every prediction in public and compare the model with margin-free closing odds in a back-test. The calibration tables on this page show predicted vs observed rates.

Related pages

18+. Gambling can be addictive; only bet what you can afford to lose. Free help: BeGambleAware.org · GamStop.