Model Performance
A probability model is judged by calibration over samples, not by single-match hits. Every engine change must move these numbers on out-of-sample matches — or it gets cut.
2026-27 Season · Cumulative (pregame predictions, no hindsight)
Matches scored
10
Model log loss
0.8502
League-prior baseline
1.0939
Edge vs baseline
+0.2437
Brier score
0.4884
Modal picks correct
6/10
Lower log loss = better. Baseline = constant league base rates (home 42%, draw 27%, away 31%). Sample is still tiny (10 matches): treat every number here as noise until at least a half-season has accumulated.
Per Gameweek
| GW | Matches | Model log loss | Baseline | Δ | Picks |
|---|---|---|---|---|---|
| GW 1 | 10 | 0.8502 | 1.0939 | +0.2437 | 6/10 |