Expected Goals, in full
A complete walk-through of the xG model as implemented here — logistic regression base with a per-coefficient penalty, 93 features, temporal train/test split, Extended Kalman Filter talent posteriors, calibration curves, and cross-site validation against MoneyPuck and Hockey-Reference. The full coefficient table, both eras, is its own ledger: the coefficients.
Overview
Expected Goals (xG) assigns a goal probability to every unblocked shot attempt (Fenwick shot) based on the shot's geometry and context — before anyone scores. The number answers: given where and how this shot was taken, what fraction of identical shots go in?
The model produces four variants per shot, not one. Each strips out a different piece of observed skill:
| Column | What it controls for | Primary use |
|---|---|---|
| xg_unadjusted | Shot geometry + context only | Closest to public xG; shot-quality baselines |
| xg_shooter_adj | Adds shooter talent posterior | Goalie evaluation (removes shooter skill) |
| xg_goalie_adj | Adds goalie talent posterior | Shooter evaluation (removes goalie skill) |
| xg_fully_adj | Both shooter + goalie posteriors | Overachievement vs a player's own established skill |
xg_unadjusted = sigmoid(base_logit + E) xg_shooter_adj = sigmoid(base_logit + E + μ_shooter) xg_goalie_adj = sigmoid(base_logit + E − μ_goalie) xg_fully_adj = sigmoid(base_logit + E + μ_shooter − μ_goalie)
Four formulas, one model. base_logit is a single fitted shot-quality regression, and μ_shooter and μ_goalie are the same two numbers from a single talent filter; the columns differ only in which of those terms is added before the sigmoid. Nothing is refit per column, which is what makes the four comparable on identical shots.
E is the league scoring environment — one number per season, learned sequentially as that season's shots arrive — and it sits under all four columns alike. The column names describe talent adjustments only; "unadjusted" means no shooter or goalie adjustment, not no environment. That matters because raw xG is shot quality for an average shooter against an average goalie in that season's environment, and average shooters do not take power-play shots — raw alone runs a few percent under on the man advantage, and the shooter-adjusted variant is what corrects for it.
The reason for four columns is that different questions need different baselines, and the rule for choosing is: whose personnel does this adjustment encode, and do they persist? Shooter adjustment on the for-side encodes your own shooters; goalie adjustment on the against-side encodes your own goalie; the opposite sides encode your opponents, which is mostly schedule. For descriptive work, include your own personnel and exclude your opponents'. For forecasting, include the personnel that persist — goalies do, skaters turn over. Evaluating a goalie requires a baseline that does not contain that goalie, which is xg_shooter_adj; evaluating a shooter requires one that does not contain that shooter, which is xg_goalie_adj. Measuring how far a player is beating their own established skill needs a baseline containing both, which is xg_fully_adj.
Where xG flows downstream
Every downstream metric in the system pulls from these columns. Aggregation rolls them up per (player, game, strength state) into the xgf_* / xga_* / ixg_* columns that appear in the leaderboard. Net Rating takes the on-ice xgf_shooter_adj and xga_shooter_adj sums — the same headline pairing the site leads with, deliberately not the fully-adjusted columns: goalie adjustment incorporates the team's own goaltending into the defensive half of a skater metric rather than neutralising it, which is not what a play-driving rating should measure. RAPM uses the same columns at shift grain. The GSAx goalie metric is xg_shooter_adj − actual_goal per shot, summed over a goalie's games — a baseline that excludes the goalie being judged.
Pipeline
The xG suite is a four-stage chain. The first two stages run offline, once per model refit; the last two run chronologically on every new game.
Every unblocked shot attempt │ │ FEATURE BUILD — geometry, game context, categorical encodings ▼ BASE MODEL FIT (offline, once per refit) │ penalised logistic regression, one ridge per coefficient block, │ fit on an earlier slice of seasons and tested on a later one, │ with the pre-shot shooter-minus-goalie talent difference entered │ as a fixed, unpenalised offset — so the coefficients stop │ absorbing the average talent of whoever happened to be shooting │ → coefficients + calibration metadata ▼ CHRONOLOGICAL SCORER (every shot, in time order) │ read each player's talent estimate and the league environment │ as they stood before the shot, score the shot, then update │ shooter, goalie and environment together with the outcome │ → four values per shot: │ unadjusted · shooter-adjusted · goalie-adjusted · fully-adjusted ▼ AGGREGATION │ rolled up per player, per game, per strength state └→ xGF / xGA / individual xG
The chronological order in stage three is the constraint that makes the rest trustworthy. Each shot's talent estimates are read before that shot is processed, and only updated afterward. No shot can be scored using information about how it, or anything after it, turned out — the model never sees the result it is predicting.
The base model is fit offline and its coefficients frozen to a stored artifact. Scoring reads that artifact; the two stages never run together, so a refit in progress cannot contaminate live numbers. Because the fitter takes the pre-shot talent difference as a fixed offset rather than leaving it for the coefficients to soak up, the base model's own strength-state intercepts carry no passenger: a power-play shot is not coded as inherently more dangerous merely because power-play shooters are better than average — that talent is credited to the shooter term, leaving the base geometry to describe the shot itself.
The fitter is a hand-rolled penalised logistic regression rather than an off-the-shelf one, for a specific reason: it applies a separate penalty strength per block of coefficients instead of one global setting. That is what lets the strength-state geometry terms be shrunk hard toward the 5v5 baseline — where the data is thin, the model should fall back on what it knows — while the main effects stay free. A single global penalty cannot express that, and at this sample size a single global penalty would in any case do essentially nothing: roughly 300,000 training shots outweigh the penalty term by the same factor, which leaves the regularisation decorative rather than real.
Rink adjustment
For most of the archive, shot coordinates were logged by each arena's own scorer, and every scorer distorts distance and lateral position a little differently. The base model reads those coordinates as given, so the distortion lands directly in xG. It is also persistent: a goalie plays half his games in one building, so he and his backups share it, and it is the kind of error a goalie rating cannot tell from talent.
The correction works arena by arena and season by season. Visiting teams' shot distances at an arena are lined up, quantile for quantile, against the league's road distances for the same season, which gives a smooth map from what that arena recorded to what the league would have recorded. Visitors are the yardstick because a team's road schedule is close to a random sample of the league, so shooting style cancels and the arena is what is left. The map is applied to every event at that arena, home and away, shots and non-shots alike, before any geometry is computed, so every feature reads one consistent coordinate system.
Scoring a game may only use what was known before it, so an arena's map is built from its own earlier games, which in October means almost nothing. Arena bias carries from one year to the next (a correlation of 0.755), so each arena-season starts from its own previous map and moves toward the current evidence as it accumulates. An arena with fewer than 5,000 visiting shots in a season is not mapped.
It stops after 2021-22, and that was measured. Through 2021-22 the gap between an arena's own team at home and on the road tracks the visitors' bias there (a slope of 0.70), the signature of a scorer mislocating everybody. From 2022-23 the slope falls to about 0.29: what is left is mostly the home team genuinely suppressing its opponents' chances, and adjusting that away would remove real hockey. The spread of the bias across arenas drops from 1.71 ft in 2021-22 to 1.21 ft in 2022-23, and on 2022-23 the adjustment made calibration worse.
What it buys is fairness across arenas, not accuracy. In 2021-22, the last season it applies to, it cut the spread between arenas from 14.31 to 7.80 percentage points while log loss did not move.
Two limits. The map that can be used in scoring, built without looking ahead, is weaker than a full-season one: it removes some of an arena's bias, not all of it. And the measurement study covers 2012-13 onward, so 2010-11 and 2011-12 are adjusted on the strength of the later seasons and were never checked directly.
Re-run policy
A refit is triggered manually when a meaningful volume of new seasons has accumulated. It produces two sets of coefficients, not one: a chip-era model for 2023-24 onward, and a pre-chip model for 2010-11 through 2022-23. Within that pre-chip span, seasons through 2021-22 have their coordinates corrected for the recording arena's own scorer before fitting; 2022-23 is scored as recorded, because by then what separates one arena from another is mostly real home-ice effect rather than scorer error. Scoring and aggregation then run automatically after each scrape cycle. A refit forces a re-score of every shot in the archive, which takes about forty minutes; an ordinary night only scores the games that were played, continuing the talent filter from where it stopped rather than starting it over.
Features
Built by xg_features.build_feature_matrix. The full design matrix runs to 93 features; the table below groups them by block rather than listing all 93 individually. Coefficient values are not reproduced here — the fitted vector lives in the stored model artifact.
| # | Feature | Category | Notes |
|---|---|---|---|
| 1 | dist_to_net | Geometric (5v5 baseline) | Euclidean feet to nearest net post (NHL coords; arena-adjusted through 2021-22) |
| 2 | shot_angle | Geometric (5v5 baseline) | Signed in x, over the full half-turn: 0 = straight-on, π/2 = on the goal line, past π/2 = behind the net, π = directly behind it (same arena adjustment as distance) |
| 3 | dist_x_angle | Geometric (5v5 baseline) | Interaction term |
| 4 | tp_* | Geometric (5v5 baseline) | Tensor product of natural-cubic-spline bases in distance and angle — 26 columns that let danger be a curved surface over the two rather than a straight line in each. The three terms above are its linear corner, so the log-linear model is nested inside it. The angle margin carries a knot past π/2 so the surface keeps bending as it wraps behind the net. |
| 5 | strength_geometry_* | Strength geometry | Per-state deviation of dist_to_net/shot_angle/dist_x_angle from the 5v5 baseline, ridged toward zero (DEVIATION_RIDGE=100) |
| 6 | strength_all_* | Strength | One dummy per skater-pair state: 5v5, 5v4, 5v3, 4v3, 4v5, 6v5, 6v4, 5v6, 4v4, 3v3, other |
| 7 | is_home | Contextual | Shooter on home team |
| 8 | is_rebound | Contextual | Prior Fenwick by same team within 3s, same period |
| 9 | own_net_empty | Contextual | Shooting team's own goalie pulled — still a live situation, since the shooter's own net being empty does not change what they are shooting at |
| 10 | crossed_royal_road | Geometric | Puck crossed the centre line between the faceoff dots within 1-3s before the shot, forcing the goalie to move laterally across the crease. The third most valuable feature in the model. |
| 11 | time_since_prev_event | Prior event | Seconds since previous on-ice play event (capped 30s) |
| 12 | dist_from_prev_event | Prior event | Feet from previous on-ice play event location (capped 200ft) |
| 13 | prev_event_same_team | Prior event | Previous event was by the shooting team |
| 14 | prev_event_faceoff | Prior event | Faceoff just won → shot from the dot |
| 15 | prev_event_hit | Prior event | Just got/gave a hit; possession compromised |
| 16 | prev_event_giveaway | Prior event | Loose-puck recovery in unusual spot |
| 17 | prev_event_takeaway | Prior event | |
| 18 | prev_event_blocked-shot | Prior event | Chaotic post-block recovery |
| 19 | prev_event_missed-shot | Prior event | |
| 20 | prev_event_shot-on-goal | Prior event | Non-rebound trailing shot |
| 21 | shot_type_wrist | Categorical | |
| 22 | shot_type_snap | Categorical | |
| 23 | shot_type_slap | Categorical | |
| 24 | shot_type_backhand | Categorical | |
| 25 | shot_type_tip-in | Categorical | Negative relative to wrist at same distance |
| 26 | shot_type_deflected | Categorical | |
| 27 | shot_type_wrap-around | Categorical | |
| 28 | shot_type_poke | Categorical | 1,300+ samples, distinct goal rate — kept separate from 'other' |
| 29 | shot_type_bat | Categorical | Same — 11–13% goal rate, high enough for own one-hot |
Skater-pair strength states
Strength is keyed on the skater pair rather than a coarse power-play/penalty-kill split: 5v5, 5v4, 5v3, 4v3, 4v5, 6v5, 6v4, 4v4, 3v3, and a catch-all other. A two-man advantage (5v3) has its own coefficient, separate from a one-man advantage (5v4). There is no 5v6 state: that bucket would be entirely empty-net shots, which are excluded from the model, so it would carry no data. own_net_empty remains — a shooter whose own goalie is pulled is still shooting at a defended net.
Hierarchical shot geometry
There is no 5v5 interaction column, so the main-effect geometry coefficients (dist_to_net, shot_angle, dist_x_angle) are 5v5 geometry. Every other strength state carries a strength_geometry_* deviation from that baseline, ridged toward zero with DEVIATION_RIDGE = 100. That ridge value came from a walk-forward sweep: held-out AUC peaked on the 100–300 plateau (0.7566) against 0.7559 for fully free per-state geometry and 0.7552 for collapsing every state onto 5v5 — collapsing is the worst option, fully free overfits the thin states. See Ablation: the per-state geometry deviations alone are worth roughly 80% of the whole strength apparatus.
Prior-event features
The “data about the last event” block: time and distance since the prior event, whether it was the shooting team's possession, and one-hots for the prior event type. They carry the speed-of-play and possession-context signal that distinguishes a clean sniper's shot from a contested point shot at the same coordinates. Candidates for “prior event” are restricted to on-ice play events — coordinate-less types like stoppage, period-end, and game-end are excluded, since a goal also generates one of those at its own timestamp and an unfiltered feed let that no-coordinate row become the goal's own predecessor (see Pipeline).
Why poke and bat get their own one-hots
Poke and bat shots get their own one-hot features rather than folding into shot_type_other: each has 1,300+ samples and an empirical goal rate of 11–13%, distinct enough that lumping them into the generic "other" bucket would inflate that baseline and bias every other shot-type coefficient toward it.
Empty-net shots are excluded from the model entirely rather than isolated with their own feature; the reasoning is in the Calibration section.
Tips and deflections at same distance
Tip-in and deflected shot types score conditional on geometry: controlled for distance and angle, a tipped shot converts less often than a clean wrist shot — it is harder to aim. The unconditional goal rate for tips is higher because tips happen closer in. The model decomposes these two effects rather than conflating them.
Base model
The refit produces two separate penalised logistic regression models — each carrying its own penalty strength per block of coefficients rather than one global setting for the whole model: a chip-era model trained on 2023-24 onward, and a pre-chip model trained on 2010-11 through 2022-23. Each shot is scored against the model from its own era. Splitting the fit isolates the coordinate-precision regime change at the 2023-10-10 chip-tracking rollout — pre-chip rink-recorded x/y were noisier, so geometry coefficients fit on chip-era data don't transfer cleanly backward.
Temporal train/test split
Each model splits temporally, not randomly. The chip-era model trains on 2023-24 + 2024-25 and tests on the in-progress 2025-26 season. The pre-chip model trains on 2010-11 through 2021-22 and tests on 2022-23. Both surface forward-in-time generalization rather than the in-sample leakage that k-fold cross-validation would allow on time-series shot data.
Diagnostics (chip-era model)
| Metric | Train | Test |
|---|---|---|
| Log-loss | 0.21798 | 0.22583 |
| Brier score | 0.05882 | 0.06099 |
| ROC AUC | 0.75964 | 0.74547 |
| ECE (10-bin) | — | 0.00527 |
Train sits at 237,826 Fenwick shots against a 6.77% goal rate; test at 116,222 shots. The train/test gap is small for log-loss (+0.008) and Brier (+0.002), indicating the model generalizes without overfitting. Test ECE of 0.00527 says predicted probabilities match empirical goal rates closely across the probability range. See the Calibration section for the per-strength breakdown. These figures are for the static chip-era fit; the rolling model that actually scores shots refits every 250 games and at every season boundary across 18 segments on a 1,300-game window, with mean absolute segment bias of 3.00%.
Why AUC is not comparable across populations. Empty-net shots convert at about 55% against roughly 6.8% for everything else — trivially easy to separate from a genuine scoring chance, which inflates AUC without reflecting how well the model reads shot quality against a goaltender. They are excluded from the model entirely, so this AUC describes only the hard part of the problem: shots a goalie actually has to stop.
The comparison that isolates the effect of that exclusion is a controlled one: fit two models on identical games, one including empty-net shots and one without, then score both on the same goalie-in-only test set. On that matched population, excluding empty-net shots does not cost accuracy — AUC moves by +0.00024, with log-loss, Brier and ECE all moving the same direction. AUC is only meaningful when the scored population is held fixed, which is also why the chip-era and pre-chip models are always reported and compared within their own era rather than against each other.
Where the danger is
Distance and angle do not enter as two straight lines. Danger is a surface over both at once: the odds collapse over the first fifteen feet and flatten past forty, and how steeply they fall depends on the angle, so a shot from thirty feet in the middle of the ice is a different shot from one thirty feet out on the goal line. The model estimates that surface with a tensor-product spline — a smooth two-dimensional basis in distance and angle, fitted jointly rather than one curve per axis — which is why the contours below bend around the crease instead of running parallel to it.
The angle is measured as a half-turn, not a quarter-turn: zero straight on, a right angle on the goal line, and beyond that, behind the net. That sounds like bookkeeping and is not. Measured as a quarter-turn — which is how this model read the ice until September 2026 — a wraparound from five feet behind the goal line and a one-timer from five feet in front of it are the same distance and the same angle, so they are the same row of the design matrix and get the same number. They are not the same shot. Over the chip era, attempts from behind the line convert at 3.9% against 11.8% for their distance-matched mirror in front. Signing the angle is what lets the surface tell them apart.
Loading…
Model structure
The model uses 93 features, 26 of them the spline surface above. There is no 5v5 interaction column, so the main-effect geometry coefficients (dist_to_net, shot_angle, dist_x_angle) are the 5v5 geometry, and every other strength state (5v4, 5v3, 4v3, 4v5, 6v5, 6v4, 5v6, 4v4, 3v3, other) carries a coefficient block for its deviation from that 5v5 baseline, ridged toward zero. Goalie status (own_net_empty, own_net_empty) is carried as its own feature rather than folded into the strength categories, so an empty net composes with whatever skater state it occurs in instead of multiplying out the category count. See Features for the full list and Ablation for what each block is worth. Every fitted value, with units, is published in the coefficients writeup.
Pre-chip model
The pre-chip model shares the same penalised fitting approach and produces coefficients with matching signs and ordering to the chip-era model — slap and snap shots rank highest, wrap-arounds and tip-ins lowest, distance and angle penalize scoring in the expected direction. Coefficients differ in magnitude where coordinate noise washes out the geometric signal:dist_to_net and shot_angle coefficients sit modestly closer to zero, because the pre-chip x/y are recorded by rink scorers rather than puck-tracking chips, and that noise dilutes the per-foot and per-radian effect. The contextual block is essentially unaffected, since those features come from event metadata, not from coordinates.
Calibration
A calibrated probability model should have predicted probability equal to empirical goal rate in each bin. The chart below plots mean predicted xG (x-axis) against actual goal rate (y-axis) for ten equal-size bins. Dots scaled by bin population. The diagonal is perfect calibration.
Loading…
Feature ablation
Each row drops one feature (or feature group) and re-scores the test set. Δ log-loss measures how much worse the model gets — larger is more important. Negative values mean the model is marginally better without the feature (noise features that the L2 penalty doesn't fully suppress).
Loading…
Bayesian skill adjustment (EKF)
The base model assigns a probability to every shot independent of who took it or who was in net. The Bayesian layer tracks each player's latent shooting or save talent as a running posterior on the logit scale, alongside a third posterior — the league scoring environment E, one number per season — then folds all three into the final xG variants.
Chronological walk
The scorer walks every shot the base model was trained on (the is_fenwick_for_xg mask) in strict chronological order from 2010-11 through today, in a single pass. Talent posteriors carry across season boundaries and across the 2023-10-10 chip-era boundary. For each shot:
- Read the current posterior (μ, σ²) for the shooter, the goalie and
E— no leakage. - Compute all four xG variants using those priors.
- Update shooter, goalie and
Etogether using the (goal − p̂) residual via the EKF step below.
EKF update equations
p = sigmoid(base_logit + E + μ_shooter − μ_goalie)
w = p · (1 − p) ← Fisher information weight
Shooter update:
σ²_new = 1 / (1/σ² + w)
μ_new = μ + σ²_new · (y − p) ← y ∈ {0, 1}
Goalie update (opposite sign — saves are good):
σ²_new = 1 / (1/σ² + w)
μ_new = μ − σ²_new · (y − p)
Environment update (same residual, its own gain):
σ²_E,new = 1 / (1/σ²_E + w)
E_new = E + σ²_E,new · (y − p)
Between shots (variance drift):
σ² ← min(σ² + σ²_walk, σ²_max) ← talent can change over timeThis is an extended Kalman filter applied to a Bernoulli observation model. The Fisher information weight w = p(1−p) is largest for 50/50 shots and smallest for near-certain outcomes (empty nets, point-blank tap-ins), which correctly downweights the talent signal from those events — and, per the gate above, empty-net goals, shootout attempts and penalty shots do not reach this update at all, since the base model never trained on them either.
Season-boundary centring
Once per season boundary, shooter and goalie talent are re-centred: chance-weighted (by raw xG, in the odds domain — a headcount mean would centre on the marginal fourth-liner rather than the shots that actually happen) over the trailing 365 days, the league's active shooters and goalies are made to net to zero. Whatever the centring removes is not discarded — it is added into E, so every probability already computed is unchanged at the instant of the re-frame. Within a season the frame is fixed, so a player's in-season trajectory reflects only their own shots, not a moving league average.
Asymmetric priors
- Shooter μ₀
- −0.05 — a little below the chance-weighted average of active shooters
- Shooter σ²₀
- 0.04 — SD ≈ 0.20 logits; tight enough to keep next-season calibration honest, wide enough for a debutant's own shots to move him
- Shooter σ²_walk
- 1 × 10⁻⁵ per shot taken
- Goalie μ₀
- −0.05 — replacement level, relative to the chance-weighted active league
- Goalie σ²₀
- 0.0025 — SD ≈ 0.05 logits; still tight, goalies are shrunk hard until they prove otherwise
- Goalie σ²_walk
- 3 × 10⁻⁷ per shot faced — treats goalie talent as near-static within a season
- σ²_max
- 0.5 — variance cap prevents variance from growing unboundedly on inactive players
The goalie prior anchors to the environment term rather than an arbitrary zero, so μ₀ = −0.05 describes replacement level relative to the chance-weighted active league. It is checked by regressing next season's realized conversion rate on this season's end-of-season posterior — the slope should run toward 1; it falls short, and the residual reads as context that does not persist (linemates, team defence) rather than an unmodeled dynamic. A sweep over prior mean, prior variance and walk, chosen on pre-chip seasons and confirmed on chip-era ones, landed on −0.05 / 0.0025 / 3 × 10⁻⁷ for goalies (−0.05 / 0.04 / 1 × 10⁻⁵ for shooters, above); tightening every knob together improved that slope, year-over-year stability and log-loss at once, because a Bayesian posterior should be smoother than the raw rate it is drawn from, not noisier.
Why shooters and goalies start in different places
The asymmetry — shooters seeded at league-average, goalies seeded at replacement-level — reflects what the position is for.
For a skater, finishing is a small slice of overall value. Most of what makes a forward or defenceman valuable lives in zone exits, defensive coverage, transition play, faceoffs, the chain of events that creates the chance — none of which the shooter-talent term tries to estimate. So when we encounter a new shooter (a debutant under 100 shots), the empirical evidence is: they convert 4–8% below the chance-weighted average of established shooters. That translated to a prior mean (μ₀ = −0.05), a little below league-average, with variance (σ²₀ = 0.04) tight enough to keep the posterior meaningful in-season, wide enough for a debutant's own shots to move him as they accumulate.
For a goalie, save skill is the position. There is no “rest of the game” to make up for poor saving — puck-handling and rebound control are second-order. So a goalie selected into NHL games is making the lineup because their save talent has already separated from the AHL pool. The population we observe is therefore not centered on league-average; it's centered slightly below the average established starter, because emergency callups and third-stringers exist disproportionately in the small-sample tail. Seeding at μ₀ = −0.05 with a tight prior (σ²₀ = 0.0025) makes those callups demonstrate they belong instead of inheriting starter-caliber estimates from no evidence. After ~2,000 shots faced the posterior has moved enough for a real goalie to separate; below that the prior dominates, which is correct — 500 shots cannot establish save talent.
The walk variance follows the same logic. Goalie talent is empirically sticky season-to-season — career save% for elite goalies barely drifts — so we treat it as near-static (σ²_walk = 3e-7) and let a goalie's end-of-season posterior carry forward. Shooter finishing is noisier and more dependent on linemates and role, so a small but nonzero walk (σ²_walk = 1e-5) lets the posterior breathe across a season.
Offseason bump at season boundaries
The per-shot walk above models intrinsic talent drift while a player is active. By itself, it gives zero extra variance during a 4-month offseason gap — a player who took 200 shots across two seasons would accumulate the same total walk variance as one who took 200 shots in one season. But real talent does drift across an offseason: linemate and team changes, age, summer training, role and system changes. So when a player's season changes between consecutive observations we apply a one-time additive bump to σ²:
if season(this shot) ≠ season(last shot for this player):
σ² ← min(σ² + OFFSEASON_BUMP[kind], σ²_max)
OFFSEASON_BUMP_SHOOTER = 1 × 10⁻³ (~100 shots-worth of within-season walk)
OFFSEASON_BUMP_GOALIE = 1 × 10⁻⁴ (~333 shots-worth)The bump is instant, not decaying — different physics from the chip-era boost below. The chip-era boost decays because it models “your prior is now stale relative to the new measurement precision”: the data got better, but the latent didn't suddenly jump, so we let σ² stay wide while we accumulate enough chip-era shots to pull the posterior onto the new precise estimate. The offseason gap is the opposite — measurement precision didn't change, but the latent itself has drifted during 4 months of no observations. That drift has already happened by the time the first October shot arrives. The clean model is an instant variance jump, then σ² shrinks naturally as Fisher information accumulates from new shots. No decay constant to tune — only the bump magnitude.
Transient catch-up boost at the chip-era boundary
The defaults above describe the steady-state walk that runs both pre-chip and chip-era. At the 2023-10-10 chip-era opener the walk variance is temporarily amplified, then exponentially decays back to default with a 30-day time constant:
walk_var(t, kind) = default[kind] + boost[kind] · exp(−Δdays / τ) where Δdays = (game_date − 2023-10-10).days, clamped ≥ 0 default[shooter] = 1 × 10⁻⁵ boost[shooter] = 1 × 10⁻³ default[goalie] = 3 × 10⁻⁷ boost[goalie] = 3 × 10⁻⁵ τ = 30 days
At the 2023-10-10 opener the shooter walk is ≈ 1.01 × 10⁻³ (~101× default); 30 days in (2023-11-09) it's ≈ 3.8 × 10⁻⁴ (~38× default); 90 days in (2024-01-08) ≈ 6.0 × 10⁻⁵ (~6× default); by mid-season 2024 the boost is negligible. The first 30–60 days of chip-era play do the catch-up — pre-chip posteriors latch onto the more precise chip-era data quickly — and the rest of the timeline operates at the same default that governed pre-chip seasons. Before 2023-10-10 the boost term is zero, so pre-chip walks are the default only.
This satisfies a specific design constraint: pre-chip career posteriors should not be discarded when chip-era play begins, but they also should not anchor the chip-era estimate when the new data is much more precise. A transient boost on the existing walk lets the same EKF math do both — the shooter carries their pre-chip career into 2023-24, then the wider walk early in chip-era play lets actual goals/saves move the posterior onto the chip-era estimate within a few weeks.
Aging drift
The offseason bump says a player's talent gets less certain over a summer; it says nothing about which way it moves. Aging drift adds the direction: at each offseason boundary the raw posterior mean shifts by drift(age), a curve fit from within-player season-over-season deltas in goals over raw xG (shooters) and raw xG over goals-against (goalies), estimated on pre-chip seasons and checked against the chip era. The shift lands before the offseason variance bump — the drift has already happened by the time the first October shot arrives, so it is applied to a posterior that is still as tight as it was in April, not to one already widened for the summer.
Goalies carry the curve at full strength; shooters carry it at half strength, because the filter already pulls a declining veteran's mean down from his own shots across a season, and the full aging correction on top of that self-correction overshoots. Both scales are checked the way the priors above are: goals over adjusted xG by age bucket should run flat, and the veteran buckets that ran high (goalies) or low (shooters) without the correction flatten toward 1.0 once the right scale is applied.
The league scoring environment
Only one thing is identifiable from a made or missed shot: the gap between the shooter's talent and the goalie's. A shot that goes in tells you something about the shooter relative to the goalie in the crease — it says nothing about whether the league, taken as a whole, is a higher- or lower-scoring place than it was last year. Fold that ambiguity into the talent terms and it has nowhere honest to go; it drifts, silently, into whichever posteriors happen to update on it. So it doesn't go there. It is booked to its own term.
We call that term E, the league scoring environment on the logit scale. It is learned one season at a time, from that season's shots only — no look-ahead, no borrowing from seasons not yet played. Shooter and goalie talent are then centred, each season, so that the league's shooters, weighted by the chances they actually take, add up to no net goals relative to average. What's left in the talent terms is talent; what moves the league's general run of scoring moves E instead.
Loading…
The puck-tracking changeover on 2023-10-10 is a break in the shot-quality model itself, not just in E — new coordinates, a new base fit either side of it. So E is only meaningful compared within an era: pre-chip seasons against each other, chip-era seasons against each other. Comparing a pre-chip E to a chip-era E compares two different rulers.
Talent posteriors
Each player's talent posterior mean evolves as the EKF processes shots chronologically. The charts below show the posterior mean on the logit scale over time for the top six shooters and top six goalies in the dataset. Values above and below zero are now literally relative to the active league: talent is re-centred chance-weighted at each season boundary so the league's active shooters and goalies net to zero, and whatever a league-wide shift the centring removes is carried instead by the environment term, not by the players. Within a season the frame is fixed, so a line's shape reflects only that player's own shots, not a moving league average.
Because the model starts from a neutral prior at season start, early-season readings are mostly prior. By mid-season the posterior has updated enough to be informative. Goalies' lines are flatter than shooters' — a consequence of the near-zero walk variance, which treats goalie talent as stable.
Loading…
x-axis is date; y-axis is posterior mean on logit scale, relative to the active league and re-centred each season. Positive = better than league-average shot quality. Goalies: positive = above-average save rate given shot quality faced. Goalie prior starts at −0.05 (replacement level).
External validation
xG is expected goals — its job is to predict future goals better than past goals do. Validation runs on two axes: internal predictive validity (does our xG outperform raw goal counts in forward-looking tests?) and cross-site agreement (does our team-season xG land in the same band as reputable public models?).
Validation is run on shots that can plausibly reach the outcome being tested: shootouts and penalty shots are excluded from training and from every figure below, and strength states are keyed on the exact skater pair. That scoping keeps every correlation on this page honest — nothing here is inflated by an outcome the model was never trying to predict.
Internal predictive validity
Three tests, run on team-game and player-season grain, offense and defense both:
- Split-half within season. For each team, partition its games randomly into halves, compute first-half xGF/60 (or xGA/60) and GF/60 (or GA/60), then correlate each with second-half goals. Repeat 50× and average. A good xG metric beats the naive goals→goals baseline.
- Year-over-year shooter talent stability. Pearson r between each shooter's posterior mean at end of season N and end of season N+1, for shooters with ≥100 shots in both seasons. Compared head-to-head with YoY stability of raw 5v5 shooting percentage. Season-boundary re-centring moves any league-wide shift into the environment term rather than the shooter posterior, so this comparison is now a cleaner read on genuine finishing-talent persistence.
- Game-level xGF → GF. Per (team, game) row at 5v5, slope and Pearson r of our xGF vs actual GF. Slope ≈ 1 with positive r indicates well-calibrated game-grain estimates.
The tables below show the three most recent seasons for readability. The same harness now runs across all sixteen seasons back to 2010-11 — pre-chip years score against the separately-fit pre-chip base model — and each table's final row pools across that full run.
Split-half, by xG variant (50 random partitions, mean Pearson r, first-half xGF → second-half GF):
| Season | xGF | xGF (sh) | xGF (sv) | xGF (sh+sv) | GF → GF (baseline) |
|---|---|---|---|---|---|
| 2023-24 | 0.531 | 0.612 | 0.515 | 0.602 | 0.421 |
| 2024-25 | 0.271 | 0.327 | 0.277 | 0.333 | 0.294 |
| 2025-26 | 0.483 | 0.589 | 0.478 | 0.580 | 0.402 |
| 16-season pooled | 0.356 | 0.477 | 0.347 | 0.470 | 0.348 |
Every xG variant beats the naive past-goals baseline in every season and in the pooled sixteen-season run — the headline confirmation that the metric does what an “expected goals” metric should. The shooter-adjusted headline metric, xGF (sh), wins split-half in two of the three seasons shown and pooled overall; in 2024-25 the fully-adjusted (sh+sv) variant nudges narrowly ahead instead. The goalie-adjusted variant runs close to the unadjusted model split-half — a touch below in most seasons, a touch above in others — which makes sense since the random partition mixes each goalie's starts across both halves rather than separating them in time.
Forward-half, by xG variant (first half of the season chronologically → second half, Pearson r):
| Season | xGF | xGF (sh) | xGF (sv) | xGF (sh+sv) | GF → GF (baseline) |
|---|---|---|---|---|---|
| 2023-24 | 0.585 | 0.572 | 0.579 | 0.565 | 0.398 |
| 2024-25 | 0.145 | 0.217 | 0.167 | 0.236 | 0.464 |
| 2025-26 | 0.437 | 0.496 | 0.458 | 0.511 | 0.344 |
| 16-season pooled | 0.333 | 0.383 | 0.334 | 0.382 | 0.292 |
The forward-half test is the honest one: it splits each team's games by calendar order rather than at random, so it mirrors how the model is actually used — shooter and goalie talents are updated only on shots that already happened, never on future ones. Read split-half as an upper bound on self-consistency and forward-half as the number that matters for anyone using this model to project the rest of a season; 2024-25 is the cautionary case, where GF → GF out-predicts every xG variant forward in time by a wide margin. Pooled across all sixteen seasons the picture is cleaner: xGF (sh) leads on offense both split-half and forward-half.
Split-half, defense (mean Pearson r, first-half xGA → second-half GA):
| Season | xGA | xGA (sh) | xGA (sv) | xGA (sh+sv) | GA → GA (baseline) |
|---|---|---|---|---|---|
| 2023-24 | 0.546 | 0.549 | 0.653 | 0.652 | 0.520 |
| 2024-25 | 0.448 | 0.444 | 0.555 | 0.549 | 0.483 |
| 2025-26 | 0.466 | 0.463 | 0.527 | 0.524 | 0.435 |
| 16-season pooled | 0.423 | 0.420 | 0.520 | 0.516 | 0.413 |
On defense the shooter-talent adjustment barely moves xGA — the players a team faces are close to a wash over a half-season — while the goalie-adjusted variant, xGA (sv), wins outright in all three seasons and pooled overall. That is the mirror image of the offense result above, and for the same reason: sv carries information (whose net it was) that a pure shot model does not have, the same way sh does on offense.
Forward-half, defense (Pearson r):
| Season | xGA | xGA (sh) | xGA (sv) | xGA (sh+sv) | GA → GA (baseline) |
|---|---|---|---|---|---|
| 2023-24 | 0.556 | 0.575 | 0.622 | 0.630 | 0.575 |
| 2024-25 | 0.371 | 0.388 | 0.458 | 0.473 | 0.521 |
| 2025-26 | 0.403 | 0.401 | 0.353 | 0.353 | 0.193 |
| 16-season pooled | 0.395 | 0.395 | 0.455 | 0.454 | 0.424 |
Forward-half defense is noisier than split-half, the same pattern as offense: 2024-25 is again a season where the raw GA → GA baseline out-predicts every variant, and 2025-26 is one where the unadjusted model edges the talent-adjusted ones. Pooled across sixteen seasons xGA (sv) still leads, which is the read that matters more than any single season's ordering. For a pooled view that adds MoneyPuck as a fourth point of comparison on both sides of the puck, see the team-level forward-half chart in “Us, against the conventional approaches” below.
YoY shooter stability (n=179, 2010-11 → 2011-12, ≥100 shots both seasons):
| Metric | Pearson r | Spearman r |
|---|---|---|
shooter_talent_mean | 0.723 | 0.697 |
| raw 5v5 sh% | 0.597 | 0.556 |
2010-11 → 2011-12 is the only season pair the generator currently computes — it's the earliest pair in the backfill, not a recent one; chip-era pairs will be added once that harness runs on chip-era seasons. The posterior is comfortably more stable year over year than raw shooting %, which is what a Bayesian estimate should be — the prior pulls hot stretches back toward league average and only sustained signal moves it.
Game-level xGF → GF (5v5, per team-game):
| Season | n | Pearson r | slope |
|---|---|---|---|
| 2023-24 | 2,800 | 0.313 | 0.709 |
| 2024-25 | 2,796 | 0.289 | 0.630 |
| 2025-26 | 2,788 | 0.303 | 0.671 |
Per-game noise dominates at this grain — that matches the public xG literature. Slopes under 1 mean game-level xGF is more dispersed than actual GF, which is what you expect when a smooth expectation is regressed against a noisy count.
xGF → GF vs public sources, by xG variant
Same team-season 5v5 grain as the cross-site agreement test, but the question here is point-estimate accuracy: of the available xG models, which one's xGF tracks actual GF most tightly? We plot Pearson r for each of our four variants alongside Hockey-Reference and MoneyPuck on the two completed seasons. (Per-game xGF isn't published by HR or MP, so the comparison aggregates to season totals.)
Loading…
Cross-site team-season agreement
Team-season xGF/xGA totals for the regular season only, using unadjusted xG for a like-for-like comparison with public shot-quality models, at 5v5 and all strengths, compared against MoneyPuck and Hockey-Reference.
Two questions, and they are different. Agreement asks whether three models looking at the same games arrive at similar team totals — high agreement means nobody is doing anything eccentric. Accuracy asks which of those totals actually tracks goals. A model can agree with everyone and still be wrong.
| 5v5 team-season | 2023-24 | 2024-25 |
|---|---|---|
| xGF agreement — ours vs MoneyPuck | 0.969 | 0.937 |
| xGF agreement — ours vs Hockey-Reference | 0.921 | 0.851 |
| xGF agreement — MoneyPuck vs Hockey-Reference | 0.929 | 0.818 |
| xGF → GF — ours (unadjusted) | 0.710 | 0.498 |
| xGF → GF — MoneyPuck | 0.715 | 0.493 |
| xGF → GF — Hockey-Reference | 0.817 | 0.457 |
The three agreement rows were computed on 13 September, before the 25 September refit, so our side of them reflects the previous model. The xGF → GF rows are recomputed with every cascade.
We agree with MoneyPuck more closely than MoneyPuck and Hockey-Reference agree with each other, which is about as much comfort as an agreement test can give. On accuracy the three of us are level: 0.710 / 0.715 / 0.817 in one season and 0.498 / 0.493 / 0.457 in the next, with the ordering changing between them. Thirty-two teams is a small sample and we would not read a gap of that size as a ranking.
Head-to-head against HockeyStats.com
A more direct test than season-total correlation: year-over-year prediction of the same skaters' next-season performance, on identical games, run against HockeyStats.com's published xG model.
| Player, 5v5 (season N → N+1) | Ours | HockeyStats | Result |
|---|---|---|---|
| Offence | 0.4671 | 0.4790 | they win, +0.012 |
| Defence | 0.2368 | 0.2309 | we win, +0.006 |
This is the most credible comparison on the page precisely because it doesn't flatter us. HockeyStats' xGF/60 edges ours slightly on predicting next season's actual goals (0.4790 to our 0.4671); on the defensive side we edge them back (0.2368 to 0.2309). Both gaps are small enough that we would not stake a claim of superiority on either one.
Head-to-head against Natural Stat Trick
A different grain than HockeyStats above: individual production and individual shot quality (ixG, shots, individual Corsi) rather than on-ice results, matched player by player for 2024-25 5v5 against a hand-delivered NST export.
| 5v5 individual, 2024-25 (n=909) | Pearson r | Levels |
|---|---|---|
| Goals | 1.0000 | mean|diff| 0.00 |
| Assists | 1.0000 | mean|diff| 0.01 |
| Shots | 1.0000 | mean|diff| 0.10 |
| Individual CF | 1.0000 | mean|diff| 0.10 |
| TOI | 0.9999 | mean|diff| 7.03 min |
| ixG | 0.9927 | totals 5,404 ours vs 5,348 theirs |
| ixG/60 | 0.9588 | means 0.501 ours vs 0.496 theirs |
Counting facts agree almost exactly across 909 matched player-seasons, so the two sides are unambiguously describing the same shifts. Individual xG tracks just as tightly — ixG/60 correlates at 0.9588, with our own season total landing about 1% above theirs.
What this isn't
We don't benchmark against private or proprietary models (Stathletes, NHL EDGE, club-internal). The HockeyStats, Natural Stat Trick, and MoneyPuck/Hockey-Reference comparisons above are the closest thing we have to a like-for-like public benchmark.
Evolving-Hockey remains out of reach without browser automation — it's a JavaScript SPA with no server-rendered tables. Natural Stat Trick's own team and game pages sit behind the same kind of Cloudflare challenge for non-browser requests, but a hand-delivered individual-player export made the head-to-head above possible.
Us, against the conventional approaches
Everything above describes what this model does. This section asks a blunter question: on the exact same games, does a reader get anything for it? Six conventional tools stand in for what most fans, media and even some analysts actually reach for — a shot-location model with no talent layer, MoneyPuck's own published xG, raw shooting and save percentage, the shot-multiplier “regress toward league average” trick, and goals-minus-xG as a finishing stat. Every one of them is measured here exactly as it would be used — same walk-forward split, same held-out seasons, same players — not quoted from someone else's writeup or flattered by picking the favorable year. Where the conventional read wins, that's printed too.
Loading…
Known limitations
- Calibration degrades with sample size, broadly. Per-strength ECE on the held-out test set is reported per strength state in the Calibration section, which renders it live from the fitted artifact rather than repeating it here. The pattern is consistent: 5v5 calibrates tightly because it has tens of thousands of test shots behind it, and every thin state is an order of magnitude worse. For the strength states that tracks sample size rather than model quality, and is probably not fixable by better features — a few hundred shots a season cannot support well-calibrated probabilities no matter what goes into the model. Shot-level probabilities in the thin states deserve real skepticism; the impact on team-season aggregates is smaller, since those states contribute few shots in total.
- Two distance bands still miss, and sample size does not explain either. They are about the same size. One is inside 10 ft: the model over-predicts 5–10 ft by about 3%. The miss is shared by rebounds and non-rebounds, sits in the same shots as the model’s top-decile over-prediction, and no distance or rebound term tested so far reaches it; it is open. The other is 30–35 ft, still under-predicted by about 5%. Beyond 22 ft the spline used to carry distance alone and missed by as much as 23%; six free-form far-distance bands fixed most of that, but the refit did not close the defencemen gap it was expected to. Defencemen’s shots still score about 2.4% more often than the model predicts (1.024, against 0.997 for forwards), so the oscillation was not its main cause. The table is in the Calibration section, rendered live from the fitted artifact.
- No goalie handedness. A glove-side vs. blocker-side effect is real — goalies have structurally different save rates on their off-hand side. Shot-type one-hots partially proxy for this (wrap-arounds skew to one side), but there is no explicit handedness feature in the current design matrix.
- No QoC adjustment in xG itself. The base model treats all shots as coming from a league-average shooter and going toward a league-average goalie, and the EKF layer adjusts for the specific individuals involved. But the base probability does not account for the quality of opponents that created the shot opportunity, and building that in would create a chicken-and-egg dependency with the talent posteriors. Net Rating used to carry a QoC/QoT adjustment of its own; it was tested and retired in 2026-09 (see the Net Rating writeup) in favor of a narrower ice-time pace adjustment on expected chances against. Competition is still adjusted for elsewhere: the forecast behind xGAx and WPAx on the Shift Performance dashboard is keyed on the terciles of competition and teammate quality a shift was actually played against. It is the shot model, and now Net Rating, that leave it alone.
- Pre-2022-23 coordinates are scorer-corrected; 2022-23 is not. The DB spans 2010-11 through today, and before puck tracking, x/y were logged by each arena's own rink scorer rather than a chip — and every scorer carries a systematic distance and lateral bias (documented in the CJ Turtoro thread), persistent enough that a goalie who plays half his games in one building would otherwise wear it as if it were talent. Seasons through 2021-22 correct for exactly that: each arena-season's visiting-team shot distribution is matched onto the league's road distribution, and the resulting map is applied to every shot at that rink, home and away alike, seeded from the arena's own prior season and shrunk toward nothing until the current season has evidence of its own. From 2022-23 the correction stops — measured directly, what is left of the arena effect by then is mostly the home team's own chance suppression rather than scorer error, and adjusting it away would remove real hockey rather than noise. What noise the correction cannot reach, and all of it for 2022-23, is left for the shooter/goalie talent posteriors to absorb as wider uncertainty bands. Chip-era posteriors remain the authoritative source for current talent estimates; pre-chip posteriors are smoother but less precise per-shot. The pre-chip base model is fit walk-forward by season — each season trained only on the pre-chip seasons before it, the same discipline the chip-era model uses — rather than one static model fit on the whole era at once and scored in-sample. 2010-11 has no prior seasons to train on, so it is a bootstrap: fit on its own first slice of games and flagged as such in the model artifact, not a genuinely out-of-sample estimate the way every later pre-chip season is. 2023-24's move to chip tracking is a separate break from the scorer-correction question above — a change in how much of the shot population gets recorded, not in where a recorded shot sits — which is why the base model is still fit and scored on either side of it.
- Shooter, goalie and league-wide talent are not separable from goals alone. If the league gets higher- or lower-scoring as a whole, goals cannot say whether that change belongs to shooters, to goalies, or to the game itself — no method can recover that split from the outcome data alone. Rather than let that ambiguity land arbitrarily on shooter or goalie posteriors — which would surface as a slow, un-anchored drift in the common level, with nothing holding it to a fixed point — the model puts a named, published number —
E, the season environment — in the place where the ambiguity actually lives, and holds shooter and goalie talent to a zero-sum, active-league frame instead.Eis only comparable within an era: 2023-24's move to chip tracking is a measurement break, not a scoring change, soEshould not be read across that boundary. - Posteriors carry across seasons. The current build runs a single chronological pass from 2010-11 to today — talent posteriors carry across both season boundaries and the chip-era boundary, with a transient walk-variance boost at 2023-10-10 that lets the precise chip-era data move posteriors quickly onto its tighter estimate before phasing back to the steady-state walk variance shared by both eras. See the Bayesian section for the boost shape and decay constants.