What does a 60% chance really mean?

By ProbaPredict Data Desk · Last reviewed

A 60% chance means that, across many similar situations, the outcome should happen about six times in ten and fail about four times. One result can't confirm or refute it. A forecaster is well calibrated when its 60% predictions come true close to 60% of the time, which is what our reliability tables check.

What does "60%" mean in practice?

Imagine 100 football matches where a model gives the home team a 60% chance of winning. If the model is well calibrated, the home team wins about 60 of them and fails to win about 40. You don't know in advance which 40, so any single one of those matches can easily go the other way. That's not the model being wrong; it's what 60% means.

Why can't one result prove a probability wrong?

Because both outcomes were possible. If you flip a coin weighted 60/40 and get tails, the coin isn't broken. The same is true for football. People often remember the 60% predictions that failed and forget the ones that came in, which makes models look worse than they are. Equally, a few wins in a row don't prove a model is good.

How do you check a probability properly?

With calibration. Group many predictions into bands (say 50–60%, 60–70%) and compare the average prediction in each band with how often those outcomes actually happened. If the 60–70% band averages 64% and happened 63% of the time, that band is well calibrated. A table of these comparisons is called a reliability table.

Our calibration tables

The first table comes from our match back-test: every home, draw and away probability the model would have given for past matches, using only data available before each matchweek. The second comes from our season back-test of title, top-places and relegation chances.

Match calibration (back-test): predicted vs observed, home/draw/away
Predicted bandOutcomesAverage predictedActually happened
0–10%9107.2%6.6%
10–20%5,35116.2%15.7%
20–30%23,13025.6%26.3%
30–40%12,19634.3%33.1%
40–50%7,65644.6%44.1%
50–60%4,17054.4%54.9%
60–70%1,88864.2%65.3%
70–80%82774.3%76.7%
80–90%26883.9%84.0%
90–100%2591.7%84.0%

Source: ProbaPredict analysis of 18,807 matches, updated Fri, 2 Oct 2026. Back-test: each match contributes three outcomes (home, draw, away), predicted using only earlier data.

Season calibration (back-test): title, top-places and relegation chances
Predicted bandTeam outcomesAverage predictedActually happened
0–5%6,0230.5%0.4%
5–15%7329.3%10.2%
15–35%63923.9%21.9%
35–65%52348.6%50.9%
65–85%22275.2%73.9%
85–95%11790.5%94.0%
95–100%23498.9%97.9%

Source: ProbaPredict season back-test of 8,490 team outcomes, updated Fri, 2 Oct 2026. Seasons re-simulated after matchdays 5, 10, 20 and 30 using only results known at the time. Overall Brier score 0.045 vs 0.101 for a naive guess.

Our live predictions are calibrated in public too, on the accuracy page, banded from 30% to 100%.

Worked example 1: a weekend of 60% picks

You see ten home teams each given 60%. Expected wins: 10 × 0.6 = 6. But the actual number varies a lot: about 25% of the time you'd see exactly 6, but you'd see 4 or fewer about 17% of the time and 8 or more about 17% of the time. A "bad weekend" of 4 out of 10 is entirely consistent with perfectly calibrated forecasts.

Worked example 2: probability to fair odds

60% as fair decimal odds is 100 ÷ 60 = 1.67 (fractional 4/6, American -150). A 10 stake at 1.67 returns 16.70 if it wins. Over 100 such bets at a perfectly fair price you would win 60 × 6.70 = 402 and lose 40 × 10 = 400: break-even, as fair odds should be. Bookmaker prices include a margin, so they are usually below 1.67 for a true 60% chance; see fair odds.

What about sharpness?

Calibration isn't everything. A forecaster who says "every team has a 33% chance" is roughly calibrated but useless. Good forecasts are calibrated and sharp: they move away from the average when the evidence supports it. The Brier score rewards both; see how accurate football predictions are.

Common mistakes with probabilities

How many predictions do you need to judge calibration?

More than most people expect. With 100 predictions at 60%, random variation alone means the observed rate could reasonably land anywhere from about 50% to 70%. With 1,000 predictions, the range narrows to roughly 57–63%. That's why our calibration tables show the number of outcomes in each band: a band with a few dozen entries is much less informative than one with thousands.

It's also why we keep every graded prediction on the record and run long back-tests. Short runs of results are mostly noise, whether they look good or bad.

Frequently asked questions

If a 60% prediction loses, was it wrong?
Not necessarily. A 60% outcome should fail 40% of the time. Only many predictions together show whether the 60% was right.
What is calibration?
How closely predicted probabilities match observed frequencies. If outcomes given 60% happen 60% of the time, the forecaster is well calibrated at that level.
What is a reliability table?
A table that groups predictions into probability bands and shows how often each band actually happened. Our tables are on this page and on the methodology and accuracy pages.
Is 60% a strong prediction?
It is a clear lean, not a near-certainty. Fair odds for 60% are 1.67, and the outcome still fails in two of every five similar cases.

Related pages

18+. Gambling can be addictive; only bet what you can afford to lose. Free help: BeGambleAware.org · GamStop.