Methods · tracking data
Filling in the untracked two-thirds
Corey Sznajder tracks a third of the season by hand. The other two-thirds are not nothing: we know who played, how much, against whom, and what the play-by-play recorded. This is how the All Three Zones page turns that into a season estimate, and how much to trust it.
Whose data, and who may see it
Every count on the page comes from Corey’s All Three Zones project: zone entries and exits, retrievals, and the passes that lead to shots, recorded game by game from video. It is licensed to Post Hockey and shown only to active patrons of his Patreon; we do not sell access to it, and nothing here is served outside that gate.
Why our 5v5 time on ice, not his
The last row matters: our exposure is 5v5 time on ice from the NHL shift charts, not Corey’s own TOI column, because the untracked games only have ours. The two agree closely but not perfectly; the gap is a difference in what counts as 5v5, and it is the same in tracked and untracked games, which is what the estimate needs.
The estimate
For a player p in a game g Corey did not track, the expected count of a metric is
mu_pg = TOI_pg × r_pN × exp(c + β · x_pg)
r_pN his rate per hour of 5v5 in SEASON N, shrunk toward the
position prior: (Y/φ + a) / (E/φ + b)
Y, E his tracked count and 5v5 hours, that season at full weight
and his other seasons at ρ^(years apart); where a context
term ships, each tracked hour counts as exp(c + β · x) hours,
so the rate is what is left once that night's context is out
φ how much one game's total varies per unit of its mean: 1 for
a count, about an eighth for the xG-created columns
x_pg what the play-by-play saw that night: usage, on-ice attempts,
zone starts, own events, rush-shaped sequences, score state,
quality of competition and teammates, opponent strength
Season estimate = tracked count + Σ mu_pg over untracked gamesThe player’s own tracked rate carries most of the information. The context term is a ridge Poisson regression fitted with player effects absorbed, so it only explains why one of his games differed from another; it cannot invent a rate for a player Corey never saw. Position groups (forwards, defence) get separate priors and coefficients.
Why a season borrows from its neighbours, and why φ is not 1
One season at a time. A rate used to be pooled flat over every season a player was tracked in — one career average wearing five different dates, which cannot show a player changing and quietly rebuilt his 2022-23 out of his 2025-26. It is now estimated for each season on its own, with his other seasons entering at ρ per year of distance. ρ is measured per column, on held-out tracked games, and the measurement settled an argument: a season standing entirely alone is worse than the old flat pooling for every column — two dozen games of tracking is too thin — but nearly every column has an optimum short of flat, around 0.6 to 0.8 for the counts and 0.9 to 0.95 for the thinner xG-created columns, which lean harder on their neighbours because a season gives them less to go on. Seven of the seventy-six measured flat and ship that way. Where a player’s year-to-year change varies with his age by more than its own standard error, an age term rides along; for most columns it does not, and nothing is applied.
φ. Splitting what we see into talent and noise needs to know how noisy one game is. For a count — entries, retrievals, shot assists — the variance is the mean, and φ is 1. The xG-created columns are not counts: they add up expected-goal values over the shots a player set up, and a sum of values wobbles about an eighth as much as a count of the same size. Treating them as counts overstated their noise roughly eightfold, which drove the prior so hard that a typical player’s own passing barely entered his own estimate. φ is measured from the data, per column, and is near 1 for the counts by construction rather than by assumption.
How the interval is built
The interval. With posterior shape A = Y/φ + a, the rate’s log-variance is about 1/A, so for the untracked sum U the variance is roughly φU + U²/A: the sampling noise in games nobody counted, plus uncertainty about the rate itself. The φ is the same one as above, so the xG-created bands are narrower than they were when those columns were priced as counts. The page shows the 80% band; the point estimate is under the cursor and is what sorting uses. A player with six tracked games gets a wide band. That is the honest answer, not a defect.
Does the play-by-play help?
The competition is a baseline that ignores context entirely: shrunk rate × TOI. Thirty percent of tracked games were hidden, by game, so every player in a hidden game was hidden together. Each metric ships the context model only where it beat the baseline on those games and its level was right; otherwise it ships the baseline and says so here.
A few percent of deviance is not a transformation. It fits an earlier finding on this data, that shot-level proxies built from the play-by-play recover only about a tenth of what the tracking knows about a shot. The player’s tracked rate does the work; the context term trims the games where he played more, or less, or against different people than usual.
How much do the tracked games tell you?
Split-half reliability of the raw per-60 rates within a player-season, tracked games divided at random, Spearman-Brown corrected to the full tracked sample. Low numbers mean the metric needs many games before it says something about the player rather than the games.
The lists are in the appendix.
Is the tracked sample representative?
Corey chooses which games to track, so the tracked third is not a random draw. On the things the play-by-play can see, the difference between a player’s tracked and untracked games is:
The context features exist partly to absorb this: if his tracked games ran hotter than a player’s season, the model sees the lower on-ice attempts and TOI in the untracked games and scales down. It cannot absorb selection on things the play-by-play does not record — which is why the table below exists.
Whom he tracked them against. The check above covers quantities the play-by-play carries for every game, so the model can already see those skews. This one covers the opposition in Corey’s own terms, which it cannot. Each row is the difference between the opponents faced in his tracked games and in the untracked ones, as a share of the untracked mean and as a fraction of the spread between teams.
The opponent table
The expected box score
The columns above count what Corey tracked. This section predicts something else entirely: the player’s real NHL box score. It exists because the two got confused once, and the confusion was ours.
The obvious move is to add up the expected goals on the shots a player set up and call the result expected assists. We did, briefly. Measured against actual 5v5 assists it came in at 0.65 for primary and 0.45 for secondary, and the band around the implied point total missed the real figure far more often than it should. The cause is not a modelling error and no modelling fixes it: Corey records a primary passer on about 74% of the shots he tracks and a secondary on about 38%, while the NHL awards a primary assist on nearly every goal. A pass nobody wrote down cannot be credited to anyone.
How the gap is measured and divided out
What does work is measuring that gap and dividing by it. Two quantities are observable for any season he tracks — the share of his shots carrying a passer in each slot, and the share of his shots we can pair to a play-by-play event. Their product explains most of the season-to-season variation, which matters because a correction fitted separately to each past season would say nothing about next season. The residual is carried in the interval rather than assumed away.
expected total = c_G · (own xG) + c_1 · (xG created, 1st pass)/(fill1 · match)
+ c_2 · (xG created, 2nd pass)/(fill2 · match)
fitted per quantity, non-negative, identity link: points really are
goals plus assists, so a model that multiplies them is harder to checkOne fit per column
Each column is fitted against its own target rather than summed from the others. Adding a calibrated xA1 to a calibrated xA2 does not give a correctly covered xA, because the two are correlated and each band is calibrated separately. Six quantities, six fits.
Validation is leave-one-season-out, because the job is to price a season the model was not fitted on. Holding each season out in turn, the league level lands between 0.92 and 1.10 and the 80% band contains the player’s real total between 77% and 85% of the time, against a nominal 80%, across all twelve quantities and five seasons (research/a3z_xpoints.json).
Why the interval is not made wider to be safe
An interval that is too wide is not the safe error. The first version covered far more than the nominal 80% because it charged twice for the same randomness, and a band that always contains the answer tells you nothing about anybody. The width is calibrated to the claim, on training seasons only.
Below five tracked games the model runs about 14% high, because the estimate feeding it falls back to a positional average. Those cells are left blank. An empty cell is honest; a number is not.
Limits
- The expected box score predicts NHL totals and is calibrated to them. The xG created columns count danger created on the passes Corey recorded, and sit about a third below real assist totals by construction. Do not read one as the other.
- Regular season, 5v5, 2021-22 onward. Earlier seasons use different workbook layouts and are not standardised yet.
- Thirty-two counts, twelve percentages and three derived counts (shot, rush and forecheck contributions, each a shot total plus its shot assists). Corey tracks more still; these are the ones with enough events per game to estimate. Percentages are ratios of two estimated counts and are shown as points, not intervals.
- Pass types split his primary shot assists by where the pass came from, so they sum to his shot-assist total rather than adding to it. Entry defence now carries the outcome of a target as well as the denial: carried in, passed in, and whether it became a chance.
- A player with no tracked games gets the position prior and a wide band, and is hidden by the default minimum of five tracked games.
- Opponent strength is the opponent’s full-season 5v5 xG rate, which includes games after the one being estimated. It is a mild look-ahead about team quality, not about the target.
- The context term is fitted to Corey’s definitions as he applies them; if his tracking changes, the model is refitted, not patched.
Appendix
The tables behind the sections above.
The full validation table
Split-half reliability and the coverage curve
Coverage. For players with at least thirty tracked games, the rate from a random 5, 10 or 20 of them was used to predict the rest. Median absolute error, as a share of the true remainder: