Model Performance
A probability model is judged by calibration over samples, not by single-match hits. Every engine change must move these numbers on out-of-sample matches — or it gets cut.
2026-27 Season · Cumulative (pregame predictions, no hindsight)
Matches scored
20
Model log loss
0.9839
League-prior baseline
1.1111
Edge vs baseline
+0.1272
Brier score
0.5782
Modal picks correct
11/20
Lower log loss = better. Baseline = constant league base rates (home 46%, draw 24%, away 30%). Sample is still tiny (20 matches): treat every number here as noise until at least a half-season has accumulated.
Per Gameweek
| GW | Matches | Model log loss | Baseline | Δ | Picks |
|---|---|---|---|---|---|
| GW 1 | 10 | 0.9003 | 0.9921 | +0.0919 | 7/10 |
| GW 2 | 10 | 1.0675 | 1.2301 | +0.1625 | 4/10 |