Methodology
A match is a probability distribution, not a guess.
1. Two numbers per team, not one
Every team carries an Attack and a Defense rating (50 = league average; one point ≈ 1.5% multiplicative change in expected goal output). The headline Power score is just their average — predictions always go back to the two underlying parameters, because a 90/50 team and a 50/90 team produce completely different scoreline distributions.
2. Process data over results
Goals are a low-frequency, high-variance outcome. Ratings are fitted on opponent-adjusted expected goals (xG / xGA) with exponential time decay (half-life ≈ 90 days), because process metrics are far more stable out-of-sample than results. Finishing over-performance regresses hard; shot creation persists.
3. Ratings to probabilities
Rating difference → expected goals for each side → full scoreline grid via a Poisson model with the Dixon-Coles low-score correction (independent Poisson systematically underprices 0-0 and 1-1). Win/draw/loss, totals and BTTS are all integrated from the same grid, so every market stays internally consistent.
4. What we do with unquantifiable factors
Nothing is allowed to move a probability "by feel". Three routes only:
- Proxy-quantify: if a factor has a data proxy (manager changes, style matchups translated into pressing / line-height features), it must prove itself by reducing out-of-sample log loss before entering the model.
- Widen, don't shift: factors with no directional evidence (dressing-room noise, new-signing integration) lower the confidence score, which shrinks output probabilities toward the league prior — they never touch expected goals.
- Market residual: we monitor the gap between model and closing-odds implied probabilities. A sudden divergence means the market knows something we don't; the match is flagged high-uncertainty.
5. Validation
The model is scored by log loss against de-margined closing odds, with strict point-in-time data (nothing the model couldn't have known before kickoff). Factors that don't improve out-of-sample calibration get cut — including fan favourites like head-to-head records.
Current ratings are a v0 preseason baseline set from public 2025-26 priors; the live pipeline (Understat / FBref process data, refit after every round) replaces them progressively. Model output is statistical information, not betting advice.