The goalie season forecast
One number per goalie, made at the end of a season and governing every shot of the next. It is a talent per shot faced, it docks goalies for age, and it runs generous on the goalies who barely played. That last part is a limit of the question, not a bug in the model.
What the number is
One row per goalie per evidence season N. The row is the forecast made at the end of N, and it governs every shot of N+1 (forecast_for_season = season + 10001, stored rather than re-derived). It is an information-weighted average of the goalie's whole history, each season discounted geometrically by how long ago it was, pulled toward a fixed anchor by a prior. A goalie with thin evidence is mostly prior, which is the point and, as the light-workload section shows, also the whole problem.
A positive number is the better goalie. It is published raw, with no display flip, and points the same way as the talent table it sits under.
It is a rate, with no volume in it. The forecast is a talent per shot faced, and nothing in the table multiplies it by a projected workload. How many shots a goalie faces depends on whether he keeps the net, and that depends on how he plays, which is the thing being forecast. Baking in a projected volume would quietly use the answer. So the consumer supplies volume: multiply the rate by the shots you expect him to face, and state your assumption about his role out loud. gsax_per_shot is the same talent as goals saved above expected per Fenwick faced at league-average shot quality. Neither column is a season total.
The formula, the units and the checks are in the appendix.
The age term
Goalies decline, and the shape of the decline is now a smooth curve rather than a straight line with a break at 30. It rises from −0.037 logits at 20 to a peak of +0.031 at 24, holds roughly level from 28 to 31 (−0.001 to −0.006), then declines to −0.040 at 35 and −0.094 at 40. It is a natural cubic spline with knots at 21, 24, 27, 31, and 35, linear outside that range, ages clipped to 20–40, and centred so a goalie of average age gets roughly zero. There is no break at 30 or anywhere else. In the table the term is applied to the output — ageing is a fact about the next season, not about what the goalie did — and stored separately (age_adjustment and forecast_logit_no_age) so any consumer can switch it off. A goalie with no birth date gets no age term.
Retired 2026-09-28. The previous version of this term was linear past 30 — AGE_BETA · max(0, age − 30) with AGE_BETA = −0.01054 — chosen because the pre-registered primary estimator cleared its threshold at that size. The spline replaces it: it scores the same out of sample and does not assert a break the data never showed.
Chosen out of sample, and it is a tie on prediction
The curve was fitted by weighted MAE against goalies who returned the following season and judged out of sample, walk-forward, with pre-season precision weights only — the published forecast is for goalies known to be playing, so its target is conditional on survival. Draft pedigree as the shrinkage anchor, a team-change term, a "late ager" interaction, a sliding delta curve, a truncated-imputation aging curve, and a joint performance-and-survival model were all tested and rejected; none beat the shipped curve; most tie it within about one standard error and a few do worse.
Walk-forward MAE among returners: 0.12996 for the new estimator against 0.12972 for the retired one — a difference of +0.00024 with a goalie-clustered standard error of 0.00084. That is a tie. The spline was not chosen for lower error; it was chosen because its shape is defensible where the old break at 30 was not, and it costs nothing on the honest metric.
The oldest goalies are still the thinnest numbers
The 34-and-up bucket is still over-docked by about 2.8% out of sample, the same as before the curve changed. The new spline does not fix the oldest goalies; the overall improvement (from 1.0121 to 0.9979) does not reach them. The decay is fixed at 0.85 per season of evidence, not fitted: a walk-forward profile from 0.7 to 1.0 was flat, and a freely fitted rate wandered between runs.
Why the age term is in the forecast, not in the walker
The obvious place for an aging effect is the talent walker itself, at the season boundary, where it already applies a goalie aging curve. That placement question predates the spline and was settled against the retired linear term: fit through 2020-21 and score what follows, and the walker's 34+ bucket already read 0.9988 on goals against over expected, and adding drift moved it away from 1.000. It has not been re-run against the spline, but the question the check answers — is there something broken at the walker's boundary right now — has not changed with the curve's shape. So talent_ekf.py and aging_curves.json are still not touched, and this stage still triggers no refit and no cascade.
Precedent, for comparison of approach
Evolving-Hockey's public GSAx, as Luke and Josh Younggren describe it, carries no aging adjustment on the goalie side at all, and McCurdy's Magnus does not either. That is a precedent for leaving an age term out, not a test of leaving it in.
The light-workload miss
The forecast is level overall and about nine per cent generous on the thinnest third of the workload: predicted goals against divided by actual reads 0.911 in the light tercile, 0.976 in the middle and 1.023 at the heavy end. That gap is the most important honest thing on this page.
These ratios were measured against the estimator retired 2026-09-28 and have not been reproduced on the decay-spline estimator shipped since. The mechanism below — a light season is light because it went badly, not because of anything the estimator did — is about workload and selection, not about which estimator produced the number, so it is kept; the exact figures should be treated as approximate until reproduced.
It is censoring, not the estimator. Classify each light season by why it was light. A season that was light by design, a backup on one roster from the first tenth of the schedule to the last, misses by 2.2%. A season that was cut short, by a demotion, a trade, an injury or a hook, misses by 9.9%.
- Light by design (a backup all season)2.2%
- Cut short (demoted, traded, hurt or pulled)9.9%
Those seasons are short because they went badly, and a forecast made before the season cannot condition on a season that never happened. The model is not so much wrong about those goalies as being asked about a set of seasons selected on their own failure.
A goalie who was pulled, demoted, traded or hurt faced fewer shots because his season went badly, and a forecast made before the season cannot know that. Expect the published number to run about 9% light on the thinnest third of workloads.
The evidence that it is censoring
It is not the estimator. Simulate the goals as draws from the model's own probabilities and recut the terciles each draw, and the light bucket comes back at 1.001. The observed 0.911 sits more than nine standard errors from that. It is not a season that starts fine and goes wrong either: on the very first appearance of the season the light tercile already reads 0.934 against 1.052 for the heavy one. Two thirds of the gap is there before a puck is faced.
Read forward instead of backward it is starker still: a goalie who continues in the league misses by 3.3%, one who is demoted by 9.4%, and a goalie whose light season was his last allowed 17% more goals than the forecast said.
The only covariate that closes the gap is a leak. Offset the forecast by the volume the goalie actually faced in the season being predicted and the tercile span collapses from 0.112 to 0.024, so 79% of the gradient is repaired, while the stayer correlation rises 67%. Improving level and prediction at the same time is what a leak looks like in this frame; nothing honest did both. It is the one quantity that would fix the light tercile and it is the outcome itself, so it is kept as a yardstick and will never ship.
What was tried and rejected, and what survives censoring
The honest candidates were run and rejected, and are recorded so nobody re-derives them. Prior workload and a short-season hazard model were both tested walk-forward; the hazard model is the better of the two and closes 0.017 of the 0.112 span for 14% of the stayer correlation. Inverse-probability weighting is not a candidate at all: reweighting changes which seasons the level check averages, not what the forecast says about any one of them.
What survives censoring is small: among by-design backups alone, the three terciles read 0.978 / 0.975 / 1.025. That five-point gradient is ordinary shrinkage, since workload is assigned on talent and a Marcel regresses everyone toward the middle, and it is a fifth the size of the headline miss. A weaker prior recovers part of it at a cost to the calibration slope, which is not a trade worth making now.
The tercile figures come from the study run on the forecast before the age term, where the gated overall level is 1.005; the shipped table, with the term, reads 1.014 on the same gate. The direction and the diagnosis are unchanged.
Where it comes from, and what checks it
The stage is scraper/compute_goalie_forecast.py and it writes one table, goalie_season_forecast. Nothing else in the pipeline reads it. It runs in the cascade immediately after compute_net_rating, so every stage in one run shares one set of season boundaries and one xG vintage, and it costs about thirty seconds.
It is a re-derivation, and the re-derivation is checked. On the same 1,531 goalie-seasons, the production stage and the research estimator agree to a maximum absolute difference of 1.4 × 10⁻⁴ logits, median 1.4 × 10⁻⁵, which is 1.2% of one materiality unit. The stage raises and writes nothing when it produces fewer than 500 rows, when any forecast is non-finite, or when the gated level ratio is missing or outside [0.95, 1.05]. The level check is gated at 800 shots faced, the same floor this project uses everywhere else.
Full methods: backend/docs/goalie-forecast.md; the age term in research/goalie_context/step1_age.md; the censoring study in research/goalie_context/lightload_level.md.
One standing caveat. The forecast consumes the talent walker's own output, shooter-adjusted xG, as its evidence. Refit the xG model, or change the walker's priors, variances, aging or centring, and every number on this page moves, because the evidence underneath it moved. The stage cascades into nothing, but plenty cascades into it: it is rerun after any cascade that touched compute_xg, and a forecast read against a stale xG is a stale forecast.
The estimator and the checks
Reference material. Each drawer opens on its own.
The estimator and its units
The estimator is the goalie's whole history on the production talent walker's own per-season evidence, each season weighted by its Fisher information and discounted geometrically by how long ago it was, all of it pulled toward a fixed anchor by a prior, plus a smooth age effect:
pred_no_age = (Σ_k d^k·I_(N-k)·o_(N-k) + P0·mu0)
------------------------------------
Σ_k d^k·I_(N-k) + P0
var = 1 / (Σ_k d^k·I_(N-k) + P0)
pred = pred_no_age + f(age)o is a season's own one-step estimate of the goalie term and I its information, summed over every season of the goalie's history with k counted in seasons elapsed — a season he missed discounts his older evidence just like a played one. The decay d is fixed at 0.85, not fitted: a walk-forward profile was flat from 0.7 to 1.0, and a freely fitted rate wandered between runs. The prior precision is P0 = 329.0099 toward an anchor of mu0 = 0.010673, re-set from a raw weighted-MAE fit of −0.046665 to put the gated level check at 1.000 in sample. There is no entrant ramp.f is the age spline described above. Each season's evidence is centred on that season's information-weighted league mean, because the walker re-centres the league at every season boundary; an uncentred cross-season comparison would read a frame move as talent.
The sign. forecast_logit is g in logit p(goal) = pre − g, where pre is the model's logit with the goalie term removed. Subtracting g lowers the goal probability, so a positive number is the better goalie. Scale: one per cent of xG(sv) is 0.01141 logits, and one tier is 0.08926.
The evidence gate and the stage's guards
The evidence gate is events.is_fenwick_for_xg, a known goalie, and a final game from the model's first season on, regular season and playoffs, which is what the walker itself consumes. A narrower regular-season mask exists and is used only for the prediction correlations, never for the evidence. The stage reads events, games and players.birth_date, and nothing else.
The research estimator reads cached frames that do not exist in production; everything the production stage does is rebuilt from the database. On identical inputs the two agree to float noise, pinned by a test at 1e-9. The residual is a vintage difference: the research cache's shooter-adjusted xG came from a different compute_xg run than the values in events today.
The stage raises and writes nothing when it produces zero rows or fewer than 500, when any forecast is non-finite, when the gated level ratio is missing or outside [0.95, 1.05], or when the row count after the write does not match what was computed. A full corpus is about 1,531 rows. It also asserts that every input was knowable before the season being forecast: no row carries a lag or a career term it could not have earned, and the season key steps forward rather than back.
An ungated per-season level target is deliberately not the stage's own check: it drifts with the league's exit rate and would read as a model regression when nothing had regressed. The excluded seasons are reported alongside rather than hidden.