A production-grade machine learning system that ingests 9+ years of football data, engineers predictive features, and generates match outcome predictions across Europe's top 5 leagues.
Complete system walkthrough: data pipeline, ML architecture, and results
What started as a data pipeline evolved into a complete prediction system. The question shifted from "can historical football data reveal patterns?" to "can we actually beat the market?"
Answering that honestly required clean data, rigorous backtesting and true out-of-sample validation. The result: a well-calibrated model that sits within ≈0.009 Brier of the bookmaker's implied probabilities using public data alone. The remaining gap is information rather than method (the market prices in team news and lineups), which is why v2 added player-level data from SofaScore.
Pull match data from FBref, Understat, and Football-Data.co.uk, plus SofaScore lineups and player stats, with rate limiting and error handling.
80 features per match: 42 match-level (rolling form, head-to-head, xG, league position, rest days) and 38 player-level from SofaScore.
Calibrated ensemble of Logistic Regression, Random Forest and XGBoost, producing home/draw/away probabilities.
Time-ordered walk-forward testing, scored with the Brier score and benchmarked against bookmaker implied probabilities.
Weekend match picks plus a bet-builder CLI pricing 16 player-prop markets with Poisson and Bernoulli models.
Scored with the Brier score (mean squared error of the home/draw/away probabilities, lower is better) and benchmarked against the sharpest line in football.
| Benchmark | Brier score |
|---|---|
| Uninformed⅓ on every outcome | 0.222 |
| This modelWalk-forward, out-of-sample | 0.198 |
| Bookmaker oddsPinnacle implied probabilities | 0.189 |
≈0.009 from the market on public data. Both scores are measured on this dataset, though not necessarily on identical matches. Next lever: lineup and team-news data.
The first v2 backtest scored 0.134 on the 2025/26 season, 36% better overnight. Results that good get investigated, not celebrated.
Two player features used post-match SofaScore ratings: information that isn't available before kick-off.
Features removed and the model retrained. The season score returned to 0.209, in line with the audited v1 baseline.
Picks are 75%+ confidence double-chance selections (win or draw) from the ensemble. Small live sample at recreational stakes: a sanity check on calibration, not a claim of long-run edge.
Comprehensive match statistics, team performance data, and historical results.
Advanced metrics including expected goals (xG), shot maps, and match events.
Historical odds from multiple bookmakers for backtesting betting strategies.
Lineups and per-player match stats (via Apify), powering the v2 player features and prop models.
True Out-of-Sample Testing
The most important thing I learned: backtesting without temporal discipline is worthless. Every result in this system comes from true out-of-sample testing. Train on data before 2023, test on 2023-2026, data the model never saw during training. And when a result looks too good, audit it: that discipline caught a post-match leak in v2 before it shipped.