Calibration

Whether a percentage from this model can be taken at face value — and whether a bigger claimed edge is a bigger real one.
Outcomes graded
4485
Bets measured
162

When the model says ~40%, what happens?

6% predicted, 23% observed, 44 calls16% predicted, 23% observed, 286 calls26% predicted, 28% observed, 2142 calls34% predicted, 33% observed, 711 calls45% predicted, 43% observed, 909 calls54% predicted, 44% observed, 250 calls64% predicted, 58% observed, 110 calls75% predicted, 52% observed, 21 calls84% predicted, 38% observed, 8 calls94% predicted, 100% observed, 4 callspredictedobserved

Each dot is one band: predicted across, observed up, sized by how many calls it holds. The dashed diagonal is a perfectly calibrated model — a dot above it means the outcome happened more often than the model said, below it less often. Distance from the line is only as meaningful as the dot is large.

Model saidNPredictedActually happenedGapReads as
0–10%445.8%22.7%+16.9ptwhen the model says about 6%, it happens 22.7% of the time
10–20%28615.8%23.4%+7.6ptwhen the model says about 16%, it happens 23.4% of the time
20–30%214226.4%27.7%+1.3ptwhen the model says about 26%, it happens 27.7% of the time
30–40%71133.6%33.3%-0.3ptwhen the model says about 34%, it happens 33.3% of the time
40–50%90945.1%43.3%-1.8ptwhen the model says about 45%, it happens 43.3% of the time
50–60%25054.3%44.4%-9.9ptwhen the model says about 54%, it happens 44.4% of the time
60–70%11064.3%58.2%-6.1ptwhen the model says about 64%, it happens 58.2% of the time
70–80%2174.6%52.4%-22.2ptonly 21 samples in this band — too few to read
80–90%883.9%37.5%-46.4ptonly 8 samples in this band — too few to read
90–100%494%100%+6ptonly 4 samples in this band — too few to read

Each finished fixture contributes three samples — home, draw and away — so a slate of 100 matches is 300 outcomes. Measured over every fixture the model priced and that has since finished — not only the ones we bet. A band with fewer than 30 samples is shown but marked unreadable: it is here because removing it would be the cherry-pick, not because it means anything yet.

CLV by model probability

posted bets only

Are we better at finding a price on favourites or on outsiders? The band is the model's own probability for the side we backed. “With a close” is the denominator for the two CLV columns — a bet whose closing line was never captured has no CLV, and counting it against the band would render a missing measurement as a worse result.

Model saidBetsAvg CLVBeat closeWith a closeWin rateW–L
under 20%2too few+0.7%50%2100%2–0
20–30%16too few+1.98%75%1662.5%10–6
30–40%11too few-0.44%54.5%1181.8%9–2
40–50%36+0.35%61.1%3658.3%21–15
50–60%14too few-0.22%35.7%1442.9%6–8
60–70%25too few-0.77%44%2560%15–10
70%+13too few-1.47%38.5%1353.8%7–6

A band marked too few holds fewer than 30 settled bets. It is shown because removing it would be the cherry-pick, not because its numbers mean anything yet.

45 bets excluded: bets outside 1X2, or on a fixture with no stored model probability — neither has a model probability to band, and inventing one would be a fabricated number.

CLV by expected value

posted bets only

Does a bigger claimed edge actually beat the close by more? If it does not, the EV number is not measuring what it says it measures — which is worth publishing whether or not it flatters us.

Claimed edgeBetsAvg CLVBeat closeWith a closeWin rateW–L
negative127+0.35%54.3%12759.1%75–52
0–3%25too few+1.15%52%2552%13–12
3–6%00
6–10%00
10%+00

A band marked too few holds fewer than 30 settled bets. It is shown because removing it would be the cherry-pick, not because its numbers mean anything yet.

10 bets excluded: bets posted before EV was recorded on the ledger.

Why two different universes: calibration is measured on every fixture because bets are SELECTED — a tip exists where the model and the market disagreed, so measuring calibration on bets alone would measure the selection rather than the model. They are not shown on one axis because a well-calibrated model can still lose money, and a profitable run can come from a badly-calibrated one. Presenting either as evidence for the other would be the easiest lie on this page to tell.