Methodology
How the rating works
A ranking is only worth arguing about if you can see how it was built. This page is the whole thing: the model, the evidence that it works, every judgement call, and the parts that are still wrong. Nothing here is a marketing claim — every number below comes from a measurement you could repeat.
The short version
Every rated set from January 2024 onward is folded into one model that solves for the skill trajectory of every player at once. When a player's true strength becomes clearer, the past is revalued — beating someone in March who turned out to be excellent is worth more than it looked at the time. The result predicts 80.0% of held-out matches correctly, and it is honest about doubt: every rating carries a deviation, and players who stop competing become progressively less certain rather than frozen in place.
What goes in
- Sets
- 1,893,177
- Games
- 4,723,774
- Events
- 16,625
- Players
- 160,527
- Weeks
- 138
- Offline singles only. Online results measure a different game — connection quality is a skill of its own — so they are excluded.
- Double elimination, 32+ entrants. Round robin and small brackets are stored but not rated: they distort strength of schedule.
- Disqualifications are discarded. A DQ is a set that was never played. Counting it as a loss punishes someone for missing a bus.
- 17,052 players actually hold a rank. Everyone else is either provisional — until a rating settles, publishing it would put a hot newcomer above a proven veteran on three good sets — or inactive. A player with no tournament in six months keeps their rating and their page but holds no rank, because a ranking is a claim about who is best now, and a year-absent player at #10 pushes ten active competitors down a place each.
Does it actually work?
The only fair test is chronological: train on the past, predict sets the model has never seen, never peek forward. Each row below is that same test after one change, scored on roughly 313,000 held-out sets. Log-loss is the number that matters — it punishes confident mistakes, which accuracy alone cannot see. A coin flip scores 0.693.
Nothing on this site was chosen because it sounded clever. Every parameter — how fast skill drifts, how much one game inside a set really tells you, how freely regional levels move — was swept and kept only if it improved this number. Several ideas that sounded good were measured and thrown away, and one parameter was deleted: the last step below removed a hand-tuned knob rather than adding one, which is the only kind of improvement worth fully trusting.
The problems we had to solve
Most of the work was not the model. It was discovering what real bracket data does to a naive one.
Japan was under-rated by about 75 points
Regions mostly play among themselves, so their rating scales can drift apart with nothing to anchor them. Japanese players were beating their predicted odds by 12% against Americans — a 6.7-sigma gap, far too large to be chance. The fix is a per-region level learned only from the sets that cross between regions, and allowed to move over time as scenes rise and fall. That gap is now 2.4%, inside statistical noise, and cross-region predictions are calibrated to within 0.1% across roughly 60,000 bridging sets.
Tournament dates lie
A set's timestamp from start.gg is when it was reported, not when it was played — organisers amend brackets weeks later. Sets from an August major were arriving dated late September, landing in the wrong week and corrupting the order of results. Rating weeks are now anchored to when the tournament actually ended.
The same person, twice on the ladder
Players change sponsors, and a new tag can look like a new human. One top player appeared under two teams simultaneously, splitting his record in half. Identities can now be merged, reversibly, and every merged history folds into one player.
Sitting out should cost certainty, not points
A player who stops competing shouldn't quietly decay, but nor should the model stay equally sure about them forever. Uncertainty now grows while a player is idle — roughly 89 points of possible drift per year — so long absences widen the deviation until the player falls off the ranked ladder naturally, and a returning veteran is treated with appropriate doubt instead of false confidence.
Momentum inside a set is real
We expected counterpicks to dominate — losing a game, then adapting. The data said the opposite, decisively: sweeps happen far more often than independent games would produce. Winning a game raises your odds of winning the next by about 1.6 times beyond what skill explains. Measured over six months of weekly refits, that effect never budged.
A 3-2 is five games, not one result
For most of this project the model read one number per set: who won. A 3-0 and a 3-2 were the same evidence, so a knob was added to soften them apart — worth +0.3% accuracy, and dishonest in a small way, because it was a fudge factor standing in for something real.
The model now reads the games. Every one of the 4,723,774 games in the corpus is its own observation, and skill is solved on a per-game scale. A set result is then no longer a separate kind of evidence — it is what the per-game odds imply once you ask who takes two out of three. That is arithmetic, not a parameter, so the knob was deleted: 65% per game is 72% per Bo3, and nothing needs fitting to connect them. The rare set recorded with no game detail still works — it is simply a blurrier look at the same quantity.
One honest complication, and it is the momentum finding above: five games from one set are not five independent facts. Counting them as if they were makes the model measurably cocky — it wins slightly more calls and pays for it in confidence, which is the trade log-loss exists to catch. So game evidence is discounted. How much was swept, not assumed, and the answer landed at 1.8 — between trusting every game fully and throwing the extra detail away. Games tell you more than the set result and less than five coin flips would.
How an event's field strength is measured
The events list labels brackets by how many ranked players actually turned up, because that is what the scene means by a stacked event and it is checkable:
- Supermajor — 30 or more of the world top 100 entered. About seven a year.
- Major — 15 or more. About twenty-five a year.
- Regional — 5 or more. About a hundred and twenty-five a year.
Two honest caveats. It counts against today's top 100, not the top 100 as it stood on the day, so an older event whose stars have since retired reads lower than it felt at the time. And it is a label on the field, not on the organisers — a beautifully run local with no ranked entrants carries no badge, which says nothing about the event except who came.
An earlier version scored strength as a multiplier from the average rating of the top ten entrants. It was replaced because it could not separate anything: the median event scored ×1.30 and 83% of events cleared ×1.20, so every real major looked identical to every strong local.
What deliberately does not affect your rating
- Event prestige. A set at a supermajor moves your rating exactly as much as the same set at a weekly. Who you beat is already the whole story; paying a bonus for the venue would double-count it.
- Placement. Ratings are built from sets, not finishes. Losing to the eventual winner in top 8 is not punished for being 9th.
- Attendance. Playing more cannot inflate a rating, only sharpen it. Grinding weeklies is not a ranking strategy.
- Character. The tier list is a read-only analysis of results. It never feeds back into anyone's rating.
- Clutch, comeback and stakes stats. These are descriptions, not inputs. Mixing them into the rating would break the one property a ranking must have: that it reflects results.
Not published yet. The tier list and the clutch / comeback / stakes stats are computed and described here, but they are held back from the site. Both are honest about their averages and still too easy to read as more precise than they are — a tier built on 200 games and one built on 40,000 do not deserve the same confidence, and the fix is a better presentation, not a bigger number. They ship when that is solved.
Where it is still weak
This is a beta, and the honest list is short but real.
- Character data is partial. Only about 44% of sets in the events crawled so far carry per-game character reporting, because most organisers only record final scores. A missing character is missing data, never an assumption — and the tier list says so.
- Regions need bridges. A scene that never travels can only be calibrated loosely against the rest of the world. Countries with few international sets carry more uncertainty than their numbers suggest.
- Small locals are invisible. The 32-entrant floor keeps strength of schedule meaningful, but it means genuinely strong players in small scenes may not appear at all yet.
- Country labels are self-reported. Roughly a quarter of players never set one, which weakens regional calibration for them.
- Names are messy. Tag changes and duplicate profiles are repaired as they are found, not automatically. If you spot yourself twice, tell us and it gets merged.
Found something wrong?
Every bug listed above was found by someone looking at a specific player and saying "that can't be right" — a phantom loss, a duplicate profile, a set dated to the wrong month. That is the most useful thing you can do here. If a number looks wrong, it may well be, and we would rather know.