Model Performance
A probability model is judged by calibration over samples, not by single-match hits. Every engine change must move these numbers on out-of-sample matches — or it gets cut.
2026-27 Season · Cumulative (pregame predictions, no hindsight)
Lower log loss = better. Baseline = constant league base rates (home 45%, draw 24%, away 31%). Sample is still tiny (10 matches): treat every number here as noise until at least a half-season has accumulated.
Per Gameweek
| GW | Matches | Model log loss | Baseline | Δ | Picks |
|---|---|---|---|---|---|
| GW 1 | 10 | 0.9244 | 0.9359 | +0.0115 | 6/10 |
Engine Calibration Benchmark (10-season backtest)
| Edge Rating engine (calibrated constants) | 0.9729 | n = 3040 |
| League-prior baseline (floor) | 1.0654 | |
| Pinnacle closing odds, de-margined (ceiling) | 0.9494 | model 0.9672 on same 2870 |
Constants (home advantage 1.2, update step 1.5, season carry 0.9, promoted prior 30.0) fitted on EPL 16/17–25/26, first 2 seasons burn-in, minimising out-of-sample 1X2 log loss. The engine clearly beats the naive prior and sits 0.018 behind the closing market — the expected picture for a goals-signal engine; the gap is the target for the xG upgrade.