Things we stopped publishing
Two metrics led this site and no longer do. Each lost a specific measured comparison, and this is the record of what those comparisons were.
Nothing on this page is a live column. xG★ and wxG are not computed for new games and are not available anywhere on the site. What follows is documentation, not a metric you can look up.
The reasoning is kept because it is reusable. The most useful thing either of them produced was a negative result: the four xG variants are close to linearly dependent, which is why a fitted blend of them mostly returns one variant and zeroes the rest. That fact still shapes how the current model is built.
xG★: what fitting the weights bought
Fixed coefficients, whether Tango's original fifth-credit or the three-term wxG blend that followed it, are a modelling choice, not a law of nature. The obvious challenge is to let the data pick the weights. xG★ was that challenge: a non-negative least-squares blend over eight features (the four xG variants and the four disjoint shot tiers: blocked, missed, saved, goal), fit separately for each strength and side to predict out-of-sample G/60.
What it bought. Out of sample it beat a fixed-coefficient metric by about 0.010 of Pearson r and plain xG(sh+sv) by 0.015. That is directionally consistent but small enough that the interval spans zero, so we could not show the advantage was real. What it cost. 64 estimated parameters, a 4v4 cell fit on essentially no data, and a standing obligation to recalibrate and rebuild the shift table whenever the underlying data changed.
It was taken out of service on 2026-09-06 for a further reason. Its weights were fit before the xG outcome-leakage fix, so they described features that no longer exist, and refitting was declined.
How it was fit
The fit ran in two stages: an all-situation prior pooled across every strength and season, then a per-cell Bayesian update anchored on the legacy weighted-shots heuristic rather than on that prior, so a cell moved off the heuristic only where its own data insisted. The eight-feature space contains both the legacy heuristic and any single xG variant as special cases, so on training data NNLS ties or beats each by construction.
The result that outlived it: the four xG variants are linearly dependent
The most interesting thing xG★ produced looks like it shouldn't be possible. The four xG variants are not merely correlated, they are linearly dependent: the identity xG(sh+sv) = xG(sh) + xG(sv) − xG(unadjusted) held at R² between 0.9998 and 0.99999 in every cell, with variance-inflation factors between 10,000 and 170,000. NNLS resolves that redundancy the only way it can: it keeps one representative of the equivalence class and pins the rest to zero, and which one it keeps is close to arbitrary. Resample a cell's rows and refit 200 times and the raw weights hop between columns. On power-play offence the goalie column came out nonzero in 51% of resamples and the unadjusted column in 95%, while the model those weights described barely moved.
That made the goalie-adjustment column read 0.000 in almost every published cell, which looked like a finding that goalie quality didn't matter. It wasn't. Reparametrized into the only combinations the redundant block actually identifies (total level, weight on the shooter adjustment, weight on the goalie adjustment), the goalie adjustment was applied at somewhere between 0.17 and 0.74 in every cell we fit, never zero. The zero in the raw table was a statement about double-counting inside a rank-deficient basis. A cell that weights xG(sh+sv) has already applied the goalie adjustment, and adding more on top would double it. It was not a statement about goaltending.
We should be careful about how far that rescues the adjustment. Refitting freely in the well-conditioned basis (the adjustment columns sit at VIF 1.1 to 2.6 instead of six figures) does not produce a clean positive goalie term: the estimates scatter from +0.77 to −2.13 across the six cells and change sign in three of them. The increments are simply small relative to the noise (at five-on-five the goalie delta has a standard deviation of 0.065 against the level's 0.465), so a free fit mostly fits noise. The defensible claim is the aggregate one: applying the full adjustment everywhere is worth +0.008 of predictive r over using unadjusted xG, with a 95% interval of [−0.005, +0.020]. Probably real, not provable, and comfortably better than letting each situation pick its own.
What happened to xG★ and wxG
Status: both retired. Neither is computed for new games or exists as a column anywhere on the site (xG★ was taken out of the API on 2026-09-06). What stands in their place is the headline pairing, xG(sh) on both sides. SAx (success above expected) is defined on it: xgf_shooter_adj − xga_shooter_adj.
xG★ is retired on the evidence in what fitting the weights bought. wxG is retired on the same standard applied to itself: scored against the plain xG(sh) inside its own blend, it loses at 5v5 offence, the largest cell on the site (0.4451 against 0.4643), and it is less repeatable than plain xG(sh) in 3 of the 4 cells (5v5 / offence, 0.828 against 0.883; 5v5 / defence, 0.624 against 0.671; pp / offence, 0.666 against 0.719). The exception is pk / defence, where xG(sh) itself is the least repeatable shot metric on the page.
The one it appeared to win
wxG's only apparent advantage was on defence, where it posted the best held-out r of the five candidates (0.29). It does not survive a target with the goalie taken out: 0.2859 becomes 0.2309, below xG(sh)'s 0.2347. The number that made wxG look like the best defensive metric was measuring next month's goaltending, not this month's skaters. The Weighted Shots page draws the two scorings side by side.
Why the goalie adjustment leaks into any defensive number
The same mechanism quietly undermines any goal-weighted or fully-adjusted defensive number, not just wxG's. The goalie adjustment in the xG model is sigmoid(base_logit − μ_goalie). That is not a neutralization: it does not strip the goalie out of the number and hand back a goalie-independent estimate of shot danger. It incorporates the goalie's own talent posterior directly into the shot value. A goaltender who has been stopping more than expected lowers every shot's xG against them; one who has been stopping less raises it. And a goaltender's talent posterior is, by construction, a persistent quantity, still mostly the same number next month. So any defensive metric built on the fully-adjusted or goalie-adjusted xG variant, or on raw goals, is carrying next month's goalie performance forward inside this month's number. It looks predictive. It is predictive. It has nothing to do with the skaters it is supposedly measuring.
xG(sh+sv) shows it plainly. Scored on raw goals it ranks second of the five (0.2749); scored goalie-neutral it ranks fifth (0.1967), behind even plain Corsi (0.1975). Nothing changed except whether the target still contains the goalie. xG(sh) does the opposite: 0.2421 on raw goals, 0.2347 goalie-neutral. It barely moves, because it never carried a goalie term against in the first place, and that makes it the best defensive number of the five once the goalie is out of the target. wxG's defensive edge is the same artifact, arriving through its goals term instead of through the xG model's adjustment.
The lesson
Every layer of sophistication here has to justify itself against what is already on the site, on its own numbers. Fitted weights did not beat fixed coefficients: xG★'s blend does not clear its own noise. wxG's blend does not beat the plain xG inside it at 5v5 offence. And wxG's one apparent win, on defence, does not survive a goalie-neutral target. What is left standing is the simplest thing on the page that still knows where the shot came from: plain xG, shooter-adjusted on both sides because shooter talent persists on both ends of the ice.
The wider argument these two lost — weighted shots against plain expected goals, and both against the simplest thing that works — is laid out in Does it predict goals? under Shot metrics.