Methods · badges
Badges
A badge is a claim that a player is among the best tenth of his position at one named act. This is how a claim is scored, what had to replicate before it was allowed in, how the badges relate to each other, and what we tried and threw out. The live list is on the Badges dashboard; the definitions are in the glossary.
The rule
Every badge uses one rule. Score the player, shrink the score toward his position group by how little his sample says, and ask how much of what is left sits above the group’s top-decile line.
component z_i = (rate_i - mu_i) / tau_i on the group prior N(mu_i, tau_i^2)
score = sum(w_i z_i) / sum(w_i) w_i = split-half reliability of component i
posterior = mu + B (score - mu), sd = sqrt(B) se, B = tau^2 / (tau^2 + se^2)
cut = mu + 1.2816 tau the 90th percentile of TALENT, not of raw scores
P = 1 - Phi((cut - posterior) / sd)
earned if post is in the top 10% of the group AND P >= 0.35
once earned, kept while P >= 0.35Level check. If the machinery is calibrated, the probabilities across a position group sum to a tenth of it. Loading…
How the spread and the errors are estimated
The spread tau is estimated with precision weights, so a tail of low-minute players does not swamp it; for rare events such as penalties drawn the plain estimate collapsed and every score inflated. Standard errors come from the spread of per-game residuals, the way Shift Performance computes them, or from a Poisson or binomial count where the source gives only totals. Spearhead and Sentry need a third method: a ratio of four season sums has no residual to take a spread of, so the error is a delete-one-game jackknife — recompute the whole season without each game in turn, and let the spread of those be the error. A player needs 20 events and ten games before he is scored at all. The three highest posterior scores among holders in each badge, position group and season wear gold, silver and bronze: a rank among holders, not a grade.
The two badges scored differently, and the three plain shares
Two badges are scored differently because their input is already a posterior. Finisher and Stopper read the talent walker, whose intervals are wider than the spread of its own means (it is tuned to forecast next season, not to be tight), so deconvolving it collapses the cut onto the average. Those two are cut at the empirical 90th percentile among qualified players instead, and so are the three badges that are plain observed shares with nothing latent behind them to deconvolve — Workhorse’s share of starts, Iron Man’s share of a schedule and Minute Muncher’s share of his bench’s ice.
The level check in full: why the band is not tight at ten percent
Level check. If the machinery is calibrated, the probabilities across a position group sum to a tenth of it. They do: the mean P sits between – and – for every shrunk badge in every season. The share who actually earn a badge is then pinned to the same tenth by the rank, and the stage that writes the table refuses to write if any badge and group with fifty or more eligible players leaves the band of 3 to 16 percent. The band is not tight at 10 because three things legitimately move it: ties at the cut, the probability floor trimming a badge whose evidence is thin, and hysteresis carrying a player who has since slipped out of the tenth. The floor was 4 percent while every live badge had cleared the reliability gate; it moved when role and outcome badges began shipping ungated, because the gate and the earned share are the same phenomenon. Across the three seasons every gated badge sits between – and – percent and every ungated one between – and –, so a floor set for the first population was the wrong floor for the second. A test holds the band so it cannot drift back.
Why the rank, and what it cost
Why the rank, and what it cost. Until 2026-09-23 a badge went to every player at better than even odds of being a top-decile talent. That is the textbook decision under symmetric loss, and it made the badges incomparable to each other. The probabilities were calibrated everywhere — mean P of about 0.10 for nearly every badge in every season — but the share clearing even odds was that tenth multiplied by how decisive the evidence was. On 2025-26 the empirically cut badges landed at 10 to 12 percent while Motor sat at 4.1, Scorer 4.7, Wall 4.8 and Forcefield 4.9. Nothing about those players was different; the components behind them replicate at 0.6 rather than 0.99, and a noisy component cannot put anyone past even odds. Ranking first and keeping the probability as a floor fixes it in one place for every badge rather than by tuning any of them: the empirically cut badges do not move at all, because their cut already was the 90th percentile, and the noisy ones rise to about a tenth. The cost is a sentence we can no longer say. A badge no longer means “more likely than not a top-decile talent”; it means “the top tenth of his group at this, with at least a one-in-three chance that is really what he is”. That is a weaker claim, and it is the honest price of a comparable one.
What had to replicate
A component enters a SKILL badge only if it replicates: split its games at random into halves, measure it in each, and the two halves must agree at 0.5 or better after the Spearman-Brown correction to full length. Role and outcome badges are not gated on it — see the three kinds below — though the number is measured and published for them just the same. The weights in every composite are these reliabilities, so they cannot be tuned to flatter anyone. There is exactly one exception, and it is marked as one: Enforcer sets fighting, hitting and size level at 1.00 each, rather than at the 0.72 and 0.97 they replicate at, because it describes a style rather than estimating a talent.
Loading…
The table behind the chart
Loading…
Three things in that table decided designs
Three things in that table decided designs. Hits replicate almost perfectly and run against offence (r about −0.5 with the for-side of xGAx among forwards), so a hustle badge built on hits would decorate players who never have the puck. Rebounds and rush shots do not replicate inside a season. And net penalties replicate worse than either half, because the players who draw penalties also take them (r 0.65). Later additions, measured the same way: competition faced relative to a player’s own team that game (0.93 forwards, 0.92 defencemen, against 0.69 and 0.55 for raw quality of competition, which carries schedule noise), competition minus teammate quality (0.98), share of the team’s penalty-kill time (0.98), and NHL EDGE bursts over 20 miles an hour per 60, whose year-to-year r is 0.88 with no correction at all.
Three kinds of claim
Three kinds, and the difference is the whole point of the house style, so it is drawn on the seal itself. A skill is an ink disc: a repeatable ability. Every component behind one has to agree with itself at 0.5 or better across random halves of a season, because a thing a player cannot do again was never a thing he could do. A role is a slate plate: how he was deployed this season, and how he played it. It is a fact about a season rather than a claim about the man, so it is not gated on replication — a coach who uses a player differently after Christmas has not made November false. An outcome is a hollow ring: what happened with him on the ice, luck and goaltending and bounces included. It is not gated either, and it is not expected to repeat; that is what makes it an outcome. The board and the player page group by those three and can be filtered to one.
A skill badge is not a compliment either. It says the ability is real and repeatable, not that it helps: Beanbag holders block shots at a rate that replicates at 0.77, and blocking runs with shots against. Repeatable and useful are different questions, and only the first one earns a disc.
None of these role badges predicts goals, and none is meant to — nor does Beanbag, the skill badge next to them. The badge is not neutral about winning. It is a description of a player, and the description is not a compliment.
Stars. A seal carries two stars when the player earned the badge the season before as well, and three when he has it three years running. Read them for what they are: the site holds three seasons, so three stars means every season there is, and it is not rare — of the badges held in 2025-26, a third carry three stars. They will mean more as the history lengthens.
Outcomes: Scorer and Ledger
Outcomes count what happened, not what he did. Scorer is goals per 60 at 5v5 and on the power play — the two states a scoring role is actually played in. It is deliberately not Finisher: Finisher reads the shooting posterior behind GAx(sh) and says who should score, and every season the two lists disagree, because a goal is a shot plus a goalie plus a bounce. Defencemen do not get a Scorer seal; their goal rate agreed with itself at 0.47 across random halves of a season, under the bar, because a dozen goals cannot be told from luck. Ledger is the on-ice goal difference, measured against what his own team managed with him on the bench — the honest version of plus-minus, and still not a skill: his goalie’s night and his shooters’ night are both in it.
Roles: how a player is used, or how he plays
Roles say how a player is used, or how he plays, and they make no claim about whether it works. The ones about use are all measured against his own bench, because deployment is one coach dividing one team’s ice and a share needs a bench to be a share of. Enforcer is the exception and the reason the sentence says “or how he plays”: it is a style rather than a deployment, and fighting, hitting, sitting in the box and being big are things a man does and is on any roster, so they are scored against the league’s position group instead.
Hard Minutes
Hard Minutes asks who a coach sends out against the other side’s best, with less help, and on the kill. All three of its components are measured against his own team and position in that game, and the middle one only became so on 2026-09-23. Before that, competition minus teammates was a raw difference — and on a bad team every player’s teammates are bad, so the whole roster scored well. It correlated −0.69 with team win share for forwards and −0.73 for defencemen, and dragged the badge to −0.54: nine earners each on Chicago and Vancouver, none on Colorado or Buffalo. Taking the team mean out leaves the question the badge always claimed to ask — which players on THIS bench draw the hard assignment — and the correlation with team quality falls to +0.09.
Minute Muncher
Minute Muncher is his share of the ice his own team gave his position group: the most repeatable number on the site at 0.99, and a statement about the coach, not the player. It was raw ice time per game until 2026-09-23, and making it a share changed nothing at all — the two correlate 1.00 within a position group in all three seasons, because a team’s forwards split about 180 minutes a night and its defencemen about 120 whatever kind of team it is. The denominator is right now, and it was never the problem.
Iron Man
Iron Man is games played over the last three seasons as a share of his team’s, and is the only badge that looks past the current one; a player who has not been in the league all three years is not scored at all, because a rookie at 82 of 82 has not made the claim yet.
Spearhead and Sentry
Spearhead and Sentry read the implied purpose of a player’s shifts — how often the bench sent him out to chase or to hold — and, since 2026-09-23, they read it against his own bench. Each is his share of his team’s attack- or defend-purpose 5v5 ice divided by his share of its 5v5 ice overall: one when the coach turns to him for that job exactly as often as he turns to him at all, above one when he reaches for him. The first version asked what share of a player’s own shifts carried the purpose, and that is a fact about his team before it is a fact about him. A team that leads a lot defends a lot; a team that trails a lot attacks a lot. Sentry’s per-team earner count correlated +0.37 to +0.55 with team win share and Spearhead’s −0.47 to −0.69, and in 2025-26 three fifths of every Spearhead in the league sat on one roster. Against the bench, the largest share on any one team falls to under one in ten, and the correlation with winning to between −0.12 and +0.37. The new number is a different one rather than a rescaled one — the two agree at 0.02 to 0.22 — and, against expectation, it also repeats better: 0.62 and 0.64 for Spearhead, 0.61 and 0.50 for Sentry, against 0.46/0.43 and 0.50/0.44 before. Taking a team out removed a large and perfectly stable term, and what was left still repeated more, which is what a real difference between the men on one bench looks like. Both were held on the old number until 2026-09-23, on a gate that was never the right test for a role badge.
The estimator that decided Spearhead and Sentry
One more thing about the aggregation, because it decided the badge. A per-game ratio divides one small noisy number by another, and a season figure built by averaging those ratios chases whichever night went strangest: measured that way the component replicates at 0.17 and 0.31 for Spearhead and 0.36 and 0.22 for Sentry, and the badge then went to between seven and fifteen forwards in a league of five hundred — one and a half to three percent, against the tenth it promises — because nothing measured that badly can clear the confidence floor. The published figure pools first and divides once, over all four season totals, so every second of ice counts the same; it replicates two to four times better. Same definition, same data, a different estimator, and it is the whole difference between a badge and noise.
Enforcer
Enforcer is a composite, not a fight count: fighting majors, hits per 60, penalty minutes per 60 and size, over every skater at his position. It had an eligibility floor of three fights until 2026-09-23, which made it a ranking of fighters and left a tenth of the twenty defencemen who fight at two men. The fight does its work through the weight now instead. Fighting, hitting and size are the three things the badge is about and they are weighted level, 27% of the score each, with penalty minutes taking the remaining 20% at their measured reliability. Those three equal weights are the only numbers in the badge system chosen rather than measured — fighting replicates at 0.72 and hitting at 0.97, and they are set equal anyway, because the badge is a description of a style rather than an estimate of a talent. The fight count enters as its square root, so the first fight is worth most, about eight tenths of a standard deviation on its own, and a tenth one cannot run away with the badge. You do not have to fight to earn it: Sammy Blais threw 26.5 hits per 60 without dropping the gloves once and is a holder. But a fighter with ordinary hits still outranks a heavy hitter who never fought — Logan Stanley fought eight times, hit 5.2 times per 60 and is fourth among defencemen — and that is the balance the square root buys.
Quick Draw
Quick Draw rates a centre’s faceoffs, and says out loud that the rating is nearly the raw number. Every draw is a contest between two men, so it is fitted as a Bradley-Terry model over the draws themselves, which adjusts for who each centre lined up against; and each draw is weighted by what it is worth. That weight is the expected-goal swing riding on the outcome over the following fifteen seconds, which is where the even-strength effect has saturated:
| Draw | xG swing, 15s |
|---|---|
| End zone, even strength | 0.0157 |
| Neutral zone, even strength | 0.0059 |
| End zone, special teams | 0.0247 |
| Neutral zone, special teams | 0.0051 |
Two things about that table are worth saying. The stake has to be symmetric — the same number whichever centre wins — because our faceoff rows record the zone and the strength from the WINNER’s point of view, so the same physical draw is an offensive-zone draw for one man and a defensive-zone draw for the other. An earlier build weighted by the winner’s side alone and produced a rating that was offensive-zone wins measured against defensive-zone losses: it correlated 0.11 with faceoff percentage and had a winger eighth. And the spread is small. An end-zone draw is worth under three neutral-zone draws, and every centre takes a similar mix, so the weighting reorders almost nobody: the finished rating correlates 0.97 with raw faceoff percentage and repeats about as well, 0.86 against 0.86. It is a better number, and it is barely a different one. The badge says so.
Beanbag
Beanbag — Target Dummy until 2026-09-23, and a role badge until the same day — is the share of the shot attempts against him that he blocks himself, at 5v5 and on the kill. It was blocks per 60, which rewarded being hemmed in: a player pinned in his own end gets more chances to block one. Dividing by the siege leaves how often he is the one who steps in front, and it moves the list. The forward who stood second on the old count blocked 5.1 a night and was stepping in front of 6.9% of the attempts against him, the lowest rate of the three; he is off the new top three, and the forward now at the head of it blocks 10.7%. It is a skill rather than a role because that rate replicates at 0.77 — it is a thing he can do, and doing it is still not the same as helping.
What the role badges do not claim, in numbers
None of these role badges predicts goals, and none is meant to — nor does Beanbag, the skill badge next to them. Blocking runs with shots against, because a player blocks shots in the end where the puck is: blocks per 60 sit at −0.32 to −0.40 with xGAx. Hitting runs against offence, about −0.5 with the for-side of xGAx among forwards, because a team hits when it does not have the puck. Fighting sits near −0.2. Minute Muncher needs a different disclaimer: ice time does track play-driving, r 0.62 among forwards and 0.30 among defencemen, but that runs through the coach’s judgement rather than through anything the badge measures, so it is his opinion with a seal on it and not a second one. Those are the numbers for the share of the bench the badge now uses, and they are the numbers for the raw minutes it used before, to two decimal places. Every live role badge is pointed at the following season’s on-ice goal share in the table below — holders against everyone else — so the disclaimers keep having to earn themselves, and Enforcer’s does more than that: forwards who hold it go on to a 44 to 46 percent on-ice goal share the following season against 50 for everyone else, at p 0.001 and 0.011. Sentry does the same thing now that it is measured against the bench: forwards who hold it go on to 42 and 44 percent against 50, at p 0.0004 and 0.0007, and Spearhead’s forwards to 53 and 52. That is the badges finding the fourth line and the first line, which is what a deployment badge should find, and it is also the clearest statement available that they are not compliments. The badge is not neutral about winning. It is a description of a player, and the description is not a compliment. They are here because professionals look at them, and because describing a player is a different job from rating one. Penalty minutes appear as one component of Enforcer and nowhere else: there is no penalty-minute badge, and there is not going to be.
How the badges relate: the crosstab
Each badge is meant to name a different act, but the set is not independent. A few links are by design: Forcefield’s only component is one of Ice Tilter’s two, and Spearhead and Sentry divide by the same share of a team’s ice. Beyond those, good players do several things, so badges travel together. The table is the correlation between badge scores.
Loading…
A study of all three public seasons, forwards and defencemen separately, found this.
- No badge is a copy of another. We set the line for “near-redundant” at an R² of 0.90 against all the other badges before looking, and none comes close. The most predictable is Minute Muncher (0.76 for forwards, 0.74 for defencemen), and for defencemen Ice Tilter (0.75), which contains Forcefield’s component by construction.
- One cluster is “the coach’s top offensive player.” A single component carries 35% of the variation for forwards and 30% for defencemen: Scorer, Spearhead, Minute Muncher, Ice Tilter, Ledger, Netfront, Finisher and Playmaker on one side, Sentry on the other. This is why stars collect several badges. It is a feature of hockey, since coaches give the most ice and the attacking shifts to the players who score, not a flaw in any one badge.
- Smaller families recur every season. Forwards: grit (Motor, Enforcer, Beanbag), puck-carrying speed (Transporter, Speedster, Playmaker leaning in) and trusted minutes (Hard Minutes, Iron Man, Minute Muncher). Defencemen: offence and puck (Finisher, Playmaker, Transporter, Speedster, Spearhead), workload (Hard Minutes, Minute Muncher, Iron Man) and defensive results (Forcefield, Ice Tilter, with Ledger and Wall leaning in).
- Some badges say something nothing else says. Quick Draw is the least related to any other badge (R² 0.15). Among the rest, Beanbag, Iron Man and Speedster for forwards, and Beanbag, Wall, Boomerang, Enforcer and Ledger for defencemen, are barely predictable from the other badges.
- The closest pairs are mostly pairs by design. Spearhead and Sentry mirror each other, and no player held both in three seasons. Outside the built-in links the closest pairs are Playmaker with Transporter and Finisher with Scorer, and even those share at most a third of their holders.
Reading the table, pair by pair
Reading it. Ice Tilter correlates with most of the rest because it is the outcome they produce, which is why it is drawn as a hollow ring and kept out of the skill set. Playmaker and Transporter move together because the player who carries the puck in is usually the one who makes the next pass. We tried removing that by scoring Playmaker on what is left after carrying. It stripped the badge from Jack Hughes and Jack Eichel and gave it to forwards who pass in the zone and rarely carry, so the overlap stays and is printed here instead. Motor runs against Finisher among forwards: forecheckers and finishers are different people. Speedster is close to unrelated to everything among forwards, Ice Tilter included: fast is not the same as good. Among defencemen it is a different story. The fast ones are the ones who carry the puck in (0.56 with Transporter), and speed has a modest tie to results (0.29).
Collectively exhaustive? Parallel analysis keeps 4 components for forwards and 3 for defencemen, so the set spans several real dimensions, not two. The rest is linemates, deployment and noise.
What this study cannot say
It uses players with enough games to be scored on every badge, which leans to regulars, though correlations over everyone tell the same story. Each badge’s own noise pulls its correlations down, so true abilities could overlap somewhat more than the table shows; for the two most predictable badges the effect is small, so “nothing is near-redundant” holds where it matters, but for the noisiest badges a low R² is partly noise. A player appears up to three times, so pooled figures are point estimates and stability is judged season by season. The three A3Z-based badges pool three tracked seasons, so adjacent seasons share tracking. Ice Tilter is still in progress and was kept in.
Do they hold up
Three tests, the same ones the site used to choose xGAx over success. Split-half on the badge score. Award the badge on each half of a season independently and count how often the two agree. And ask whether holding a badge says anything about the next season: is it re-earned, and does the outcome it should predict differ between holders and the rest.
Loading…
The full table
Loading…
What we tried and dropped
Tap any line for what happened.
An 80% line, and then a 50% one.
The plan asked for four-in-five confidence. The scores were calibrated, and at 80% a badge named one to five players in a hundred: a component that replicates at 0.6 cannot place anyone in the top tenth that firmly inside a season. The line went to 50%, more likely than not — and that had the same defect one step smaller, so a probability line no longer decides anything on its own. The rank decides; 0.35 of probability is the floor under it.
A floor on shrinkage.
Players once needed their own evidence to carry half the estimate. Brett Pesce led the league in entry-denial rate over three tracked seasons, stood at 96%, and was refused Wall at 0.49. The probability already prices thin evidence. The floor went.
Watch and carried states.
“Watch” marked anyone between 35% and 50%, which read as nearly there for players a third of the way. “Carried” held last season’s badge hollow until the new season spoke; once the floor went it applied to one player a year. A badge now belongs to a season, and earlier seasons are history on the player’s page.
Spearhead and Sentry as a share of a player’s own shifts.
That is a fact about his team wearing his name: a team that leads defends, a team that trails attacks, and the badges lined up with the standings — Sentry’s per-team earner count at +0.37 to +0.55 against team win share, Spearhead’s at −0.47 to −0.69, and three fifths of the league’s Spearheads on one roster in 2025-26. Both now ask what share of the team’s purpose ice he got against his share of its ice overall. No team holds more than one earner in ten.
Averaging per-game ratios.
The first build of that fix aggregated the season as a TOI-weighted mean of the per-game ratios, which sounds like the same thing and is not: dividing one small noisy number by another gives a heavy tail, and the average follows it. The component replicated at 0.17 to 0.36 and the badge reached about two percent of forwards instead of ten. Pooling the four season totals and dividing once gets 0.50 to 0.64 and the tenth it promises. The estimator was the whole difference, on identical data.
Raw quality of competition for Hard Minutes.
It missed Nico Hischier at the 86th percentile. Measured against his own team, with teammate quality and penalty-kill share, he is 94th and holds it in all three seasons, and the badge became stable enough to return for defencemen.
Finisher for defencemen on talent alone.
It re-earned and predicted nothing (r 0.33 with next season’s goals per 60). Talent plus the shot value he produces himself predicts at 0.49, and its top tenth scores twice the rest a year later.
Forcefield against a coarse forecast.
Without the competition and teammate terciles it went to sheltered third pairs. It now forecasts against the same fine cell Shift Performance uses.
Imputed season estimates as badge inputs.
They are already shrunk, so their spread understates talent; with them the entry-denial rate showed no signal at all. Badges read raw tracked counts, pooled over three seasons because a third of one season is about 120 attempts against a regular defenceman.
Netfront as production alone.
The first version counted expected goals from within 15 feet and gave Timo Meier gold while he converted at the 20th percentile. Conversion cannot be its own component: at the net front it replicates at 0.17 in a season and 0.25 over three, and the fitted spread of true finishing there is about ten percent. So the badge is now one number, net-front goals per 60: expected goals on shots within 15 feet, tips, deflections and rebounds, times the player’s own goals-to-expected ratio shrunk toward the league’s by 87 expected goals’ worth of prior. Meier fell from first to the edge of the top tenth; Hyman, Lee and Vilardi lead.
Agitator, and penalties as a skating proxy.
Net penalties replicate at 0.43. And bursts over 20 miles an hour correlate 0.13 with penalties drawn among forwards and 0.04 among defencemen, so Speedster takes nothing from the box score. Bursts over 22 replicate well and are too rare to be a unit: a median forward has two a season.
Sources and rights
Playmaker, Transporter, Wall, Boomerang and the second halves of Netfront and Motor rest on Corey Sznajder’s All Three Zones tracking, shown under an opaque licence: a badge and its probability are public, and no count, rate, score or posterior that could be turned back into his numbers leaves the server. His figures stay with his subscribers on the All Three Zones page. Speedster reads the league’s public EDGE tracking, one request per skater-season from 2021-22.
On the ideas: shot assists as a better guide than points come from Ryan Stimson’s passing project at Hockey Graphs; possession exits and the retrieval that precedes them from Alex Novet’s transition work there; the defenceman compass, and the observation that denials are the most repeatable part of entry defence, from Sznajder himself. Knodell’s EDGE series at Puck Over The Glass found that burst rates repeat at team level and add little to shot share, which our player-level numbers bear out. Micah Blake McCurdy’s isolated-impact charts and Luke and Josh Younggren’s player cards at Evolving-Hockey set the standard for saying what a number is and is not; neither labels players, and the restraint is instructive. Outside the sport, the scarcity rule is Michelin’s, and the warning about badge inflation is every loyalty programme’s.