Back to plain xG
Tom Tango proposed weighted shots in 2014, a blunt reweighting of outcomes. We tested every layer built on top of shooter-adjusted xG since: fitted weights, a blended metric, a goal premium. None beats it, and the one place one seems to is a goaltending artifact. What survives is the plainest thing in the stack: expected goals, with the shooter adjusted for and the goalie kept out.
What each layer buys
We ran a held-out test. For players the models had never seen, the first 41 games of a season predict the rest of that season's on-ice goal rate. Five candidates: raw shot attempts (CF/CA), Tom Tango's 2014 weighted shots, xG(sh), xG(sh+sv), and wxG, a fitted blend that is now retired.
- On offence at 5v5, xG(sh) wins outright: r = 0.46, against Tango's 0.39 and Corsi's 0.35. It is the clearest result on this page.
- On defence at 5v5, scored against raw future goals, it does not. wxG has the best number (0.29). Tango (0.27) and xG(sh+sv) (0.27) also beat xG(sh) (0.24).
That last result is the one to distrust. Switch the target and watch the ranking.
- CF/CA0.23
- Tango 20140.27
- xG(sh)0.24
- xG(sh+sv)0.27
- wxG0.29 (best)
xG(sh) barely moves (0.24 to 0.23). Every other metric falls by 0.03 to 0.08, and xG(sh+sv) falls furthest. That fits the mechanism: xG(sh) keeps the goalie out. The goalie-adjusted variants carry a goaltender's talent estimate, which is persistent, so on raw goals they partly predict next month by knowing who is in net. Take the goalie out of the target and their apparent edge goes. That is the case for keeping the goalie out of xG(sh), and for retiring wxG.
A second test asks whether a metric agrees with itself: split each player-season into odd and even games and correlate the halves. Raw goals are the least repeatable number in every cell, which is the strongest argument against a goal premium. On offence xG(sh) is the most repeatable shot metric (0.88 at 5v5, 0.72 on the power play). On defence it is the least repeatable of Corsi, Tango and xG(sh+sv) (0.67 at 5v5, 0.42 on the penalty kill), a real trade for predicting better.
We are not alone in this. Knodell's fifteen-season test found scoring chances competitive with fancier models, with defence showing the least separation between any of them. Sprigings and Toumi showed in 2015 that expected goals beat Corsi by carrying shooting talent, and a 2025 skill-adjusted model reports the same kind of gain from shooter and goaltender skill.
One caveat. Everything here scores one task, forecasting a skater's future on-ice goal rate, and that is the task least kind to xG, whose job is to describe chances created whether or not they went in. A metric can be valuable and forecast no better.
Predictivity vs sample size
The same test, drawn over sample size. Every metric is measured over the first N games of a season and the target over the games after them, so nothing on the x-axis can see its own answer. Curves rise because more games means less noise, and flatten where a metric runs out of signal.
Three things to read off it. Corsi, Fenwick, shots and goals separate in that order, with goals the noisiest at every sample size, which is why nobody projects from goals alone. On offence at 41 games, Tango's weighted shots (0.39) sit above every raw rate (Corsi 0.35) and below the xG variants (xG(sh) 0.46): one fixed weight recovers a real share of what a shot-quality model buys, and it stays on the chart because it is the line any quality model has to clear. And defence is harder to predict than offence throughout, because one player drives the other team's chances far less directly than his own.
What the chart's controls do
Who switches between skaters and teams. Predicting chooses the target (GF/60, GA/60 or GF%), and every line moves with it because each metric is taken on the matching side. Strength picks the bucket. Metric switches between Pearson r, which is rank quality only and blind to a constant bias, and MAE, which is not. Show isolates the shot rates, the weighted-shots benchmark, or the four xG variants.
Read the team curves with more caution than the player curves: there are about thirty-two teams a season against nine hundred skaters, so every team point rests on a fraction of the sample. The caption under the chart reports the smallest sample behind the view you are looking at.
Limitations and open questions
Tap any line for the fuller version.
The ceiling is low, and that is the question's doing. Even the best cell, 5v5 offence at r = 0.46, explains about 0.22 of the variance in a skater's future on-ice goal rate.
Every defensive and special-teams cell does worse. On-ice goal rates are noisy and team-contaminated, so at this signal level no amount of feature engineering separates cleanly from any other. If the number is to move, it moves through a better target or more data, not better weighting.
The xG rows are a little optimistic. The talent walker's settings were tuned with the test seasons in view.
Its variances were re-tuned over all seasons, including the seasons the held-out test uses, and nobody has measured by how much. Corsi and Tango have nothing fitted. The numbers also predate the 2026-09 xG refit and will be re-measured after it.
The shot model reads the power play worse. Ranking shots by danger is harder with a man advantage.
Location and shot type carry most of the signal we have, and on the power play they carry less of it. A point shot tipped in traffic and a point shot into shins look nearly identical in the data, and the things that tell them apart (seam passes, screens, one-timers off cross-ice feeds) are not in the play-by-play. That is a gap in the model independent of anything above.
On the penalty kill, shot quality and shot volume cannot be told apart. xG(sh) scores 0.2063 against Corsi's 0.2098, with an interval that spans zero.
Scored with the goalie removed, it is a three-way tie between Tango (0.2086), xG(sh) (0.2083) and Corsi (0.2067). We do not claim shot quality beats shot volume on the penalty kill, and we do not claim it loses. This cell cannot currently tell the two apart.
Shot quality is less stable than shot volume on defence. xG(sh) repeats at 0.67 at 5v5 and 0.42 on the penalty kill, against Corsi's 0.75 and 0.55.
That is a real trade. At 5v5 it predicts better than Corsi does but wobbles more from one half-season to the next, because it is a sum over continuous per-shot values rather than a count. For ranking players over a full season the predictivity matters more. Over ten games, treat the defensive numbers with more caution than the offensive ones.
A skater's xGF is partly his linemates' finishing. Shooter adjustment is applied per shot, to whoever took it.
On-ice xGF aggregates a whole line, so a finisher who outperforms the base model lifts the on-ice xGF of every teammate who was out there, including a playmaker who touched the puck and never shot it. Some of what reads as shot creation here is a linemate riding along.
The rolling xG model relearns its own coordinate-to-goal mapping. That keeps it calibrated to the tracking system, and also bakes the system's errors in.
The base model is refit on a rolling window of recent shots, so it reflects the current relationship between reported shot location and goals. Any systematic error in how this era's tracking reports location becomes signal, not something corrected out. xG here is the probability of a goal given reported location, which in this era is not guaranteed to be true shot danger.
Everything here rests on the underlying xG model. This page does not validate it.
xG(sh) is built entirely from that model's outputs, so if the base model is miscalibrated xG(sh) inherits the error in full. The xG writeup checks the model against other public xG models and states its own caveats, which apply here too.
Three seasons, about 110 test players per cell. Most comparisons on this page carry an interval that spans zero.
The 5v5 offence result is the one exception. We have tried to say “we could not show a difference” rather than “there is no difference,” but the distinction is easy to lose and worth holding onto.
The tables and the history
Reference material. Each drawer opens on its own.
The held-out tables, the goalie fork and repeatability
Held-out Pearson r against the rest-of-season on-ice goal rate, 5v5. These numbers are hand-carried and predate the 2026-09 xG refit; they are re-measured after it. The xG rows carry a small optimism: the talent walker's variances were tuned over all seasons, including these test seasons, and nobody has measured by how much. CF/CA and Tango have nothing fitted.
| Metric | 5v5 offence | 5v5 defence |
|---|---|---|
| CF/CA | 0.3484 | 0.2279 |
| Tango 2014 | 0.3920 | 0.2721 |
| xG(sh) | 0.4643 | 0.2421 |
| xG(sh+sv) | 0.4635 | 0.2749 |
| wxG | 0.4451 | 0.2859 |
Offence, xG(sh) minus Tango: +0.0726, 95% CI +0.0421 to +0.1019 (bootstrap). On defence the old page printed a gap of −0.0296 (CI −0.0651 to +0.0047) beside wxG; it matches xG(sh) against Tango (0.2421 − 0.2721 = −0.0300), not wxG (−0.0438), so it is not quoted here until it is re-derived.
The goalie fork
| 5v5 defence target | CF/CA | Tango 2014 | xG(sh) | xG(sh+sv) | wxG |
|---|---|---|---|---|---|
| Raw goals | 0.2279 | 0.2721 | 0.2421 | 0.2749 | 0.2859 |
| Goalie removed | 0.1975 | 0.2199 | 0.2347 | 0.1967 | 0.2309 |
On the penalty kill, against raw goals: Corsi 0.2098, Tango 0.2150, xG(sh) 0.2063. Goalie removed: Tango 0.2086, xG(sh) 0.2083, Corsi 0.2067, a three-way tie. xG(sh) minus Corsi on raw goals: −0.0039, 95% CI −0.0374 to +0.0308, P = 0.41.
Split-half repeatability
| Cell | CF/CA | Goals | Tango 2014 | xG(sh) | xG(sh+sv) | wxG |
|---|---|---|---|---|---|---|
| 5v5 / offence | 0.872 | 0.554 | 0.861 | 0.883 | 0.880 | 0.828 |
| 5v5 / defence | 0.754 | 0.373 | 0.739 | 0.671 | 0.732 | 0.624 |
| pp / offence | 0.604 | 0.525 | 0.613 | 0.719 | 0.713 | 0.666 |
| pk / defence | 0.550 | 0.345 | 0.533 | 0.422 | 0.459 | 0.456 |
Odd against even games within a player-season, Spearman-Brown corrected.
A brief history of weighted shots
Public hockey analytics has spent twenty years arguing about how much credit to give a shot. The argument runs in one direction — toward weighting more, counting less.
It started with possession proxies that valued every attempt the same. Corsi (Tim Barnes, writing as "Vic Ferrari," circa 2007) summed every shot attempt at the net — goals, shots-on-goal, missed shots, and blocked shots — as a proxy for territorial play. Fenwick (Matt Fenwick, 2007) dropped the blocked shots, on the argument that a blocked attempt isn't the same kind of event as one that actually reaches the net.
The next move was to weight the surviving attempts by quality. Dangerous Fenwick (Dawson Sprigings, writing as "DTMAboutHeart" on Hockey-Graphs) used shot location and type to weight each unblocked attempt by its conversion likelihood — an early public attempt at what would become expected goals (xG). War-on-ice's scoring chances and Eric Tulsky's shot-quality work pushed in the same direction. By the mid-2010s, MoneyPuck, Evolving-Hockey, Hockeyviz, and Natural Stat Trick were each shipping their own xG variants — different models, same idea: replace shot count with shot probability.
Running alongside that, a different idea. In 2014 Tom Tango proposed weighted shots — not a quality model at all, but a blunt reweighting of outcomes: goals, plus one-fifth of a credit for every other attempt. His argument was purely predictive. Regressed across seasons, non-goal attempts foreshadow future goals at roughly a fifth the rate that goals themselves do, so that is what they should be worth. Puck++ later added score and venue adjustments and found the result beat score-adjusted Corsi out of sample. No shot location, no shot quality, no player adjustment — and it worked.
That left a residual question: what about the goals themselves? An xG model gives a 0.30-xG shot the same credit whether it goes in or not — that's the whole point of expected. But once you're building a single shot-driving stat for player evaluation, it's tempting to add goals back as their own term, on the theory that finishing is real and persistent. wxG, the metric this page is about, takes that theory seriously: a small attempt credit, plus xG, plus a goal premium, built on a talent-adjusted shot-quality model rather than Tango's flat weighting.
wxG and Tango's weighted shots are different metrics with different arguments behind them, and the naming is worth separating clearly: his gives every non-goal attempt a flat fifth of a goal and knows nothing about where the puck came from; wxG puts the bulk of its weight on a talent-adjusted xG model and gives non-goal attempts about a thirtieth.
What follows is the comparison that decides between them: wxG and xG★, a fitted blend of the four xG variants, against the plain xG model that sits inside both. What survives on the site today is the plainest thing on this page — not for lack of alternatives, but because it is the one that wins the comparison.
Two metrics came out of this line of work and are documented, not published: xG★, a fitted blend of the xG variants, and wxG, a weighted blend built on top of it. Both are covered in the Archive, along with the evidence against them.