What ball-by-ball cricket predicts beyond the scoreboard
A leakage-audited model of T20 cricket, built on 4,748,382 T20 and ODI deliveries. Gradient boosting on match state cuts log-loss on win probability 29% below the base rate (0.490 against 0.693) on held-out matches, and the study measured what player identity adds before building on it: 0.31%, under the bar.
| Innings | Over | Score | Probability |
|---|---|---|---|
| 1 | 0 | 0/0 | 45.7% |
| 1 | 1 | 7/0 | 50.0% |
| 1 | 2 | 12/0 | 45.7% |
| 1 | 3 | 27/0 | 60.2% |
| 1 | 4 | 51/0 | 66.0% |
| 1 | 5 | 72/0 | 77.5% |
| 1 | 6 | 92/0 | 86.5% |
| 1 | 7 | 98/0 | 85.7% |
| 1 | 8 | 110/1 | 86.5% |
| 1 | 9 | 115/1 | 84.9% |
| 1 | 10 | 127/1 | 84.9% |
| 1 | 11 | 137/1 | 85.7% |
| 1 | 12 | 161/1 | 92.6% |
| 1 | 13 | 171/1 | 91.0% |
| 1 | 14 | 191/1 | 94.7% |
| 1 | 15 | 203/1 | 93.3% |
| 1 | 16 | 204/4 | 89.8% |
| 1 | 17 | 211/4 | 89.8% |
| 1 | 18 | 220/4 | 89.8% |
| 1 | 19 | 231/5 | 87.6% |
| 2 | 0 | 0/0 | 96.9% |
| 2 | 1 | 4/0 | 96.9% |
| 2 | 2 | 25/0 | 77.5% |
| 2 | 3 | 32/1 | 92.6% |
| 2 | 4 | 36/2 | 97.9% |
| 2 | 5 | 47/3 | 99.6% |
| 2 | 6 | 52/3 | 99.9% |
| 2 | 7 | 68/3 | 99.6% |
| 2 | 8 | 72/4 | 99.9% |
| 2 | 9 | 83/5 | 100.0% |
| 2 | 10 | 88/5 | 100.0% |
| 2 | 11 | 103/5 | 99.9% |
| 2 | 12 | 119/5 | 100.0% |
| 2 | 13 | 128/6 | 100.0% |
| 2 | 14 | 134/6 | 100.0% |
| 2 | 15 | 139/6 | 100.0% |
| 2 | 16 | 143/8 | 100.0% |
| 2 | 17 | 152/8 | 100.0% |
| 2 | 18 | 154/9 | 100.0% |
The question
Given the state of a T20 match — score, wickets, balls left, the target, who is at the crease and how long they have been there — how well can you price the next ball and the result, and how much more do you gain by modelling the players themselves? The second half is the expensive decision: a hierarchical model with a parameter per batter and bowler is easy to propose and costly to maintain, so the study measured its value before building it.
It is a measurement problem more than a modelling one, and measurement problems fail quietly: a leak or an optimistic interval produces a number that looks like a discovery. So most of the engineering went into making the measurement trustworthy.
The model
A four-rung ladder, each rung scored against the one below it: a marginal baseline, then models of increasing capacity, up to gradient boosting on 27 whitelisted features of the match state. Every rung is fit on the training period only, tuned and calibrated on a later validation period — isotonic regression for win probability, temperature scaling for the next ball — and scored once on a held-out test period that no fitting decision ever saw.
On win probability the top rung reaches a test log-loss of 0.490 against 0.693 for the base rate — 29% lower, over 1,489 held-out matches. On the far harder next-ball task, eleven outcome classes, it is 4.17% better than the baseline. Fig. 1 is that model, unchanged, scoring a match from the test period.
How it was validated
Leakage discipline in a research repository is usually a convention, and conventions decay under deadline. Here it is a property of the code. The feature builder’s first statement is a select against a frozen whitelist, so an outcome column is unreachable by construction, and a test proves it by asserting byte-equality between a clean frame and one with label-bearing columns injected. The data loader raises on any read of the test split unless the leaderboard runner passes an explicit flag, so the single test touch is enforced rather than remembered.
The leakage canaries have to prove their own power before they are believed. The shuffled-identity test first requires that a planted identity signal on a synthetic split is detected, and only then that shuffling destroys it — a canary that cannot detect anything passes every leak. A poisoned-column canary and a temporal-split check run in continuous integration alongside it.
The confidence intervals resample matches, not balls. Balls inside a match are not independent, so resampling them manufactures precision; the match is the unit, and the resampling is written as a multinomial count matrix, which turns it into linear algebra rather than a loop. A test plants a between-match spread and requires the paired interval to come out more than ten times tighter than the unpaired one.
The corpus is an exact replay of 22,211 Cricsheet files through a pure transition function, pinned by hash and rebuilt byte-identically. A file that cannot be replayed is quarantined with a coded reason, never dropped — 5,457 of them, mostly formats outside the study’s scope — and 11 golden fixtures pin one real pathology each: a super over, a revised target, penalty runs, a concussion replacement.
| Predictor | Improvement over the level below |
|---|---|
| match state | +4.17 percent |
| player identity | +0.31 percent |
| per-match latent | +0.024 percent |
| justifies further work | 1.0 percent and above |
What the measurement decided
Identity is worth 0.31% in log-likelihood and a per-match latent for pitch and conditions 0.024%, both below the materiality bar, so the per-player model was not built: the measurement fell short of the bar that would justify it, rather than instinct deciding.
The evaluation protocol was fixed before the test split was read; the verdict bands were entered with the evaluation itself. The register below draws that line exactly.
- the four-rung ladder and validation-only tuning
- the metric contract and the calibration policy
- the leakage canaries, including the ladder-inversion stop rule
- the corpus and label hash pins
- the single-test-touch discipline
- the materiality bands that classify the result
- the challenger gate thresholds
- Cricsheet ball-by-ball, snapshot of 2 July 2026, pinned by hash
- Python · polars · scikit-learn · scipy · uv · GitHub Actions