Model Performance
A probability model is judged by calibration over samples, not by single-match hits. Every engine change must move these numbers on out-of-sample matches — or it gets cut.
2026-27 Season · Cumulative (pregame predictions, no hindsight)
Matches scored
9
Model log loss
1.0742
League-prior baseline
1.1888
Edge vs baseline
+0.1145
Brier score
0.6377
Modal picks correct
4/9
Lower log loss = better. Baseline = constant league base rates (home 45%, draw 26%, away 28%). Sample is still tiny (9 matches): treat every number here as noise until at least a half-season has accumulated.
Per Gameweek
| GW | Matches | Model log loss | Baseline | Δ | Picks |
|---|---|---|---|---|---|
| GW 1 | 9 | 1.0742 | 1.1888 | +0.1145 | 4/9 |