Rating Systems of the World — a primer on how they're built and what they believe
Status: Active · Last reviewed: 2026-08-11
Owner-commissioned educational deep-dive: how each region's/organization's rating system is constructed, the philosophy behind it, and how TwoMore's ELO-based tier system compares. Companion to rating-interop-model.md (the schema design) and rating-intl-audit.md (the gap audit). Sources: the 2026-08-11 three-thread research pass (primary federation regulations, API docs, algorithm papers) + origin-story verification.
The fundamental split: two families, one hybrid
Every system in the world belongs to one of two philosophical families, or deliberately straddles them:
Family A — measured estimators. "Your rating is a statistical estimate of your skill, updated by evidence." Elo (chess, 1960) is the ancestor; Glicko (1995) added the crucial second number — uncertainty — so the system knows how much to trust itself. UTR, WTN, DTB LK, and Playtomic live here. The philosophy: skill is a latent quantity; matches are noisy measurements; the algorithm's job is inference. Ratings move continuously, in both directions, and the system is honest about confidence.
Family B — earned credentials. "Your standing is what you have won, and it gates where you may compete." The KR division ladders (신인부→오픈부, 개나리→국화, KASTA groups) are the purest examples: no formula estimates you — your placement history licenses you into strata, wins ratchet you up (often permanently — 개나리 graduation never releases), and the design goal is fair brackets and anti-sandbagging, not measurement. FFT's classement is Family B with Family A aspirations: a points pyramid, but losses never subtract (a ratchet) and the ladder's real job for a century has been tournament seeding and aspiration ("monter en série").
The hybrid — NTRP. The most instructive system for us (see "How ours compares").
System by system
Elo → Glicko: the ancestors (chess, 1960 / 1995)
Arpad Elo's insight: model each player's skill as a distribution, treat every game as evidence, move both players' estimates by the gap between expected and actual outcome. Glickman's refinement: carry a rating deviation (RD) so a rating built on 3 games and one built on 300 are different claims — and let RD grow with inactivity, so the system knows when it's out of date. Philosophy: epistemic humility, formalized. Every serious modern system (WTN explicitly, UTR effectively, ours literally — player_ratings.rd) is downstream of this.
USTA NTRP — the fair-league classifier (USA, 1978)
- Origin/philosophy: built by USTA/USPTA to classify players "for more compatible matches" — the league-play machine for the world's largest recreational tennis market. NTRP does not want to measure you precisely; it wants your Tuesday-night 3.5 league to be fun and fair.
- Construction: a public coarse scale (1.5–7.0 in 0.5 bands, anchored by written "General Characteristics") sitting on a hidden two-decimal dynamic rating recalculated after every match but never published. One number per player across all formats. Entry via self-rate questionnaire; integrity via a human governance layer — auto-granted up-appeals, committee down-appeals, and the three-strike auto-disqualification that polices sandbagging (a named cultural problem).
- Blind spots: coarse bands frustrate band-edge players; one-number design can't speak to singles/doubles gaps; sectional pools drift apart (a 4.0 in one region ≠ another). The strike system exists because any system whose bands gate entry creates an incentive to stay low — the deepest lesson NTRP offers us.
FFT classement — the aspiration pyramid (France, ~1920s)
- Origin/philosophy: among the oldest continuously running rating traditions in sport, built for tournament seeding in a country where the classement is part of tennis identity. It is an ordinal ladder (NC → 40 → 30 → 15 → 5/6 … 0 → séries), climbed by beating people above you. Losses gate bonuses but never subtract — the ratchet reflects the philosophy: the classement is a career, not a snapshot.
- Construction: points per win scaled by the opponent's échelon gap, minimum win counts per échelon, monthly rolling recompute since Oct 2022 (a genuine epoch break from the historical annual reset).
- Blind spot/strength: non-numeric, direction-confusing to outsiders, and slow downward — but its competitiveness is legendary: UTR's founder explicitly cited FFT-style level-based events as producing far more competitive matches than US junior tennis, and built UTR to generalize that.
UTR — the market disruptor (private, 2008)
- Origin/philosophy: Dave Howell, a Virginia pro, wanted one global, up-to-date scale where any two players anywhere could be compared — directly inspired by FFT's classement but rebuilt as a Family-A estimator. Its killer app was college recruiting (coaches comparing a Serbian junior to a Californian one), and its product creed is "level-based play": competitive matches are more fun, so measure precisely and match tightly.
- Construction: 1.00–16.50 at two decimals; rolling 30-match/12-month window; per-match rating driven by opponent gap AND games-won margin (score margin matters — a philosophical break from win/loss purism); separate singles/doubles numbers; verified vs unverified tiers; daily recompute.
- Blind spots: recency-heavy (an injury layoff moves you); margin-sensitivity creates run-up-the-score incentives; privately owned — the scale's governance is a company's product decisions.
ITF WTN — the federation's answer (global, 2021)
- Origin/philosophy: the ITF plus national federations, mission stated plainly at launch: break down "one of sport's key barriers to participation — uneven match-ups." WTN is a participation-growth instrument wearing a rating's clothes: get every player a number so every player can find a fair match, worldwide, under federation governance (a strategic response to UTR's private land-grab).
- Construction: 40 (beginner) → 1 (elite) — deliberately reversed so beginners "count down" as they improve; separate singles and doubles numbers (explicitly not comparable); Glicko-2 engine scoring at set level; a public confidence percentage (RD in a friendly costume) with a 70% "verified" threshold; free via federations; public GraphQL API. Adopted outright by the LTA (replacing its old ratings), the ITA (US college), JTA (Japan, in progress).
- Blind spots: algorithm revisions reposition players in place (the scale itself isn't stable); coverage skews to federation-registered players — thin in exactly the casual populations apps like ours serve.
DTB Leistungsklasse — the club-league utility (Germany, ~2007)
- Origin/philosophy: grew out of regional systems (Saarland/Rheinland-Pfalz) and was nationalized to solve a specific, very German problem: Medenspiele team lineups must be ordered by strength, so strength must be a single administrable number. The LK is infrastructure, not identity — and it deliberately incentivizes activity.
- Construction: 25 (beginner) → 1 (best); one number across ALL formats (singles, doubles, league, tournaments feed the same LK); per-win improvement = points × age factor / hurdle, where the hurdle grows steeply toward LK 1 (physics-like resistance at the top); and the famous monthly decay cron — every inactive player drifts +0.1 toward 25 on the last Wednesday of each month. Play or sink.
- Blind spots: no discipline split at all (your doubles-carried LK follows you into singles brackets); no confidence concept — the number is simply administratively true.
KR ladders (KTA/KATO/KASTA/테니스타운) — credential gates (Korea)
- Philosophy: documented in full in our own research — these are licensing tiers with expiry, built for bracket fairness and anti-sandbagging in a tournament culture with formal elite (선출) segregation. Standing = qualification history; promotion mechanics are one-way doors (개나리 graduation, 챌린저 exclusion); "ratings" in the statistical sense don't exist — the KTA pair-handicap point tables (10.0→1.0) are the closest thing, and they're licensing math, not inference.
- Why it matters to us: this is the vocabulary our KR users actually speak, which is why our credential rows cite it — but it is categorically Family B, and mapping it to any Family-A number is calibration across philosophies, only defensible as labeled bands.
Japan — the fragmented market (JTA/JOP + 草トー + WTN incoming)
No unified amateur system: JOP ranking serves the competitive registered class; the huge casual 草トーナメント scene runs on per-organizer labels (オープン/A/B/C級, platform-specific ladders like テニスベア's Lv.1–9) that are explicitly non-standardized; JTA is adopting WTN top-down. Philosophy vacuum at the amateur middle — which is an opportunity for an app with its own credible internal rating.
Playtomic — the product-led rating (padel, global)
0.0–7.0 Elo-style level with a visible reliability score, seeded by questionnaire, tuned per country. Philosophy: the rating exists to make the product work (auto-matched competitive games), not to be an institution. The closest structural cousin to what a club app's internal rating actually is — and evidence that a self-owned rating can become a de facto standard if the product wins.
How ours compares
Our architecture, named honestly, is: a Family-A engine wearing a Family-B display, in a Family-B market.
| Axis | Ours | Closest kin | Notes |
|---|---|---|---|
| Engine | ELO with rd per track | WTN (Glicko-2) | Same lineage; we keep uncertainty like WTN, unlike LK/FFT |
| Public shape | 5 coarse tiers (브론즈→그랜드마스터) over a hidden precise rating | NTRP exactly | NTRP proved this structure for 47 years: precise engine, coarse honest display, less band-anxiety |
| Discipline topology | three fully independent tracks (단식/복식/혼복) | WTN/UTR (split S/D) — we go further | The most purist position in the market; nobody else runs an independent mixed track. US mixed pair-sum stays an eligibility rule, which our interop model handles as a declaration |
| Vocabulary layer | Per-market credential/scale labels (KR divisions today) | UTR-app side-by-side pattern | Tiers never are NTRP/divisions — they're annotated with them |
| Governance | Auto only (no strike/appeal layer yet) | NTRP's committee layer is the warning | The moment tiers gate session entry (they already can), sandbagging incentives exist — NTRP's strike system is the precedent to study when we see it |
| Decay | None yet (rd grows implicitly stale) | LK's cron vs Glicko's RD-inflation | A deliberate open decision: we currently neither decay ratings (LK-style) nor formally inflate rd on inactivity (Glicko-style). Worth deciding before leaderboards matter |
| Seeding | seed_rating + self-declared credentials | Playtomic questionnaire; NTRP self-rate | Same pattern as everyone; our verified-credential path (elite RPC, external ratings) is the NTRP "computer rating" analog |
The three positions we hold that the research validates:
- Hidden-precise + public-coarse (NTRP's proven structure) — band-edge honesty without false precision.
- Uncertainty as a first-class citizen (Glicko/WTN lineage) —
rdalready in the schema. - Independent tracks — stricter than anyone; defensible because consolidation is exactly where NTRP's one-number and LK's all-formats designs generate their known complaints.
The three warnings the world's systems give us:
- Sandbagging follows gating (NTRP): when tier bands gate entry, some players want low ratings. Our reliability gates + host authority are the current defense; a formal review/strike mechanism is the eventual one.
- Decay is a decision, not a default (LK vs Glicko): choose activity-incentive decay, uncertainty inflation, or neither — explicitly.
- Scales you anchor to will move under you (WTN revisions, FFT 2022): which is why the interop model versions every crosswalk and stamps every observation.
What "side by side" means concretely
A profile shows: our tier (earned, internal, the only thing we compute) · our band expressed in the market's primary scale (labeled inference — "복식 NTRP 3.5~4.0 수준" / someday "WTN ~25 Zone") · the user's verified external ratings (observations with provenance and dates — "UTR 5.2 (verified, 2026-05)") · credentials (Family-B facts — "오픈부 입상"). Four different epistemic objects, four different visual treatments, never merged into one number. That's the whole philosophy in one sentence: we compute one thing, translate honestly, display the rest as what it is.