Skip to content

Ilia DudaCo-op Jan 2027

What ball-by-ball cricket predicts beyond the scoreboard

Independent research · July 2026

A leakage-audited model of T20 cricket, built on 4,748,382 T20 and ODI deliveries. Gradient boosting on match state cuts log-loss on win probability 29% below the base rate (0.490 against 0.693) on held-out matches, and the study measured what player identity adds before building on it: 0.31%, under the bar.


Fig. 1
India v New Zealand, replayed ball by ballprobability the side batting first wins · gradient boosting on match state · calibrated on the season after its training · a match it never saw · log-odds scale
Probability India wins, before each ballWin probability for India, the side batting first, before each of the 249 balls of the final. It opens at 46%, India make 255 for 5, and by the start of the chase the model gives them 97%. The lowest it goes is 45%. New Zealand are all out for 159; India won by 96 runs.

P(India win)
> 99.9%
Innings
2 · New Zealand chasing
Before ball
over 18.6
Score
New Zealand 159/9
Needs
97 from 7 balls
This ball
wicket
Result
India won by 96 runs

Drag across the chart, or use the slider or the arrow keys, to move ball by ball · Home and End jump · the line is the model’s output before each ball

The published model, trained on matches up to November 2024 and calibrated on the season after, scoring the 2026 ICC Men’s T20 World Cup final from the held-out test period — chosen by rule, not by curve. Across all 1,489 held-out matches its log-loss is 29% lower than the base rate’s.
Probability India wins, at the start of every over
InningsOverScoreProbability
100/045.7%
117/050.0%
1212/045.7%
1327/060.2%
1451/066.0%
1572/077.5%
1692/086.5%
1798/085.7%
18110/186.5%
19115/184.9%
110127/184.9%
111137/185.7%
112161/192.6%
113171/191.0%
114191/194.7%
115203/193.3%
116204/489.8%
117211/489.8%
118220/489.8%
119231/587.6%
200/096.9%
214/096.9%
2225/077.5%
2332/192.6%
2436/297.9%
2547/399.6%
2652/399.9%
2768/399.6%
2872/499.9%
2983/5100.0%
21088/5100.0%
211103/599.9%
212119/5100.0%
213128/6100.0%
214134/6100.0%
215139/6100.0%
216143/8100.0%
217152/8100.0%
218154/9100.0%

The question

Given the state of a T20 match — score, wickets, balls left, the target, who is at the crease and how long they have been there — how well can you price the next ball and the result, and how much more do you gain by modelling the players themselves? The second half is the expensive decision: a hierarchical model with a parameter per batter and bowler is easy to propose and costly to maintain, so the study measured its value before building it.

It is a measurement problem more than a modelling one, and measurement problems fail quietly: a leak or an optimistic interval produces a number that looks like a discovery. So most of the engineering went into making the measurement trustworthy.

The model

A four-rung ladder, each rung scored against the one below it: a marginal baseline, then models of increasing capacity, up to gradient boosting on 27 whitelisted features of the match state. Every rung is fit on the training period only, tuned and calibrated on a later validation period — isotonic regression for win probability, temperature scaling for the next ball — and scored once on a held-out test period that no fitting decision ever saw.

On win probability the top rung reaches a test log-loss of 0.490 against 0.693 for the base rate — 29% lower, over 1,489 held-out matches. On the far harder next-ball task, eleven outcome classes, it is 4.17% better than the baseline. Fig. 1 is that model, unchanged, scoring a match from the test period.

How it was validated

Leakage discipline in a research repository is usually a convention, and conventions decay under deadline. Here it is a property of the code. The feature builder’s first statement is a select against a frozen whitelist, so an outcome column is unreachable by construction, and a test proves it by asserting byte-equality between a clean frame and one with label-bearing columns injected. The data loader raises on any read of the test split unless the leaderboard runner passes an explicit flag, so the single test touch is enforced rather than remembered.

The leakage canaries have to prove their own power before they are believed. The shuffled-identity test first requires that a planted identity signal on a synthetic split is detected, and only then that shuffling destroys it — a canary that cannot detect anything passes every leak. A poisoned-column canary and a temporal-split check run in continuous integration alongside it.

The confidence intervals resample matches, not balls. Balls inside a match are not independent, so resampling them manufactures precision; the match is the unit, and the resampling is written as a multinomial count matrix, which turns it into linear algebra rather than a loop. A test plants a between-match spread and requires the paired interval to come out more than ten times tighter than the unpaired one.

The corpus is an exact replay of 22,211 Cricsheet files through a pure transition function, pinned by hash and rebuilt byte-identically. A file that cannot be replayed is quarantined with a coded reason, never dropped — 5,457 of them, mostly formats outside the study’s scope — and 11 golden fixtures pin one real pathology each: a super over, a revised target, penalty runs, a concussion replacement.

Fig. 2
What each enrichment adds beyond match staterelative NLL improvement over the level below · next ball, T20 · calibrated
What each enrichment adds beyond match stateRelative improvement in negative log-likelihood over the level below, T1/T20, after calibration. Match state improves on a marginal baseline by 4.17 percent. Over match state, player identity adds 0.31 percent and a per-match latent adds 0.024 percent, both below the 1.0% threshold at which the decision rule would justify building the model, and player identity falls inside the ambiguous band between 0.3% and 1.0%.match state +4.17%player identity +0.31%per-match latent +0.024%0%1%2%3%4%5%ambiguous 0.3% to 1.0%1.0% justifies further work
Match state carries the signal. Player identity adds a real, reproducible 0.31%, but under the 1% bar for further work.
Relative improvement in negative log-likelihood over the level below (match state over a marginal baseline; player identity over match state; the per-match latent on validation), T1/T20, after calibration.
PredictorImprovement over the level below
match state+4.17 percent
player identity+0.31 percent
per-match latent+0.024 percent
justifies further work1.0 percent and above

What the measurement decided

Identity is worth 0.31% in log-likelihood and a per-match latent for pitch and conditions 0.024%, both below the materiality bar, so the per-player model was not built: the measurement fell short of the bar that would justify it, rather than instinct deciding.

The evaluation protocol was fixed before the test split was read; the verdict bands were entered with the evaluation itself. The register below draws that line exactly.

Fixed before any test split was read

  • the four-rung ladder and validation-only tuning
  • the metric contract and the calibration policy
  • the leakage canaries, including the ladder-inversion stop rule
  • the corpus and label hash pins
  • the single-test-touch discipline

Entered with the evaluation

  • the materiality bands that classify the result
  • the challenger gate thresholds

data
Cricsheet ball-by-ball, snapshot of 2 July 2026, pinned by hash
stack
Python · polars · scikit-learn · scipy · uv · GitHub Actions