Skip to content

Rating Systems of the World — a primer on how they're built and what they believe

Status: Active · Last reviewed: 2026-08-11

Owner-commissioned educational deep-dive: how each region's/organization's rating system is constructed, the philosophy behind it, and how TwoMore's ELO-based tier system compares. Companion to rating-interop-model.md (the schema design) and rating-intl-audit.md (the gap audit). Sources: the 2026-08-11 three-thread research pass (primary federation regulations, API docs, algorithm papers) + origin-story verification.

The fundamental split: two families, one hybrid

Every system in the world belongs to one of two philosophical families, or deliberately straddles them:

Family A — measured estimators. "Your rating is a statistical estimate of your skill, updated by evidence." Elo (chess, 1960) is the ancestor; Glicko (1995) added the crucial second number — uncertainty — so the system knows how much to trust itself. UTR, WTN, DTB LK, and Playtomic live here. The philosophy: skill is a latent quantity; matches are noisy measurements; the algorithm's job is inference. Ratings move continuously, in both directions, and the system is honest about confidence.

Family B — earned credentials. "Your standing is what you have won, and it gates where you may compete." The KR division ladders (신인부→오픈부, 개나리→국화, KASTA groups) are the purest examples: no formula estimates you — your placement history licenses you into strata, wins ratchet you up (often permanently — 개나리 graduation never releases), and the design goal is fair brackets and anti-sandbagging, not measurement. FFT's classement is Family B with Family A aspirations: a points pyramid, but losses never subtract (a ratchet) and the ladder's real job for a century has been tournament seeding and aspiration ("monter en série").

The hybrid — NTRP. The most instructive system for us (see "How ours compares").

System by system

Elo → Glicko: the ancestors (chess, 1960 / 1995)

Arpad Elo's insight: model each player's skill as a distribution, treat every game as evidence, move both players' estimates by the gap between expected and actual outcome. Glickman's refinement: carry a rating deviation (RD) so a rating built on 3 games and one built on 300 are different claims — and let RD grow with inactivity, so the system knows when it's out of date. Philosophy: epistemic humility, formalized. Every serious modern system (WTN explicitly, UTR effectively, ours literally — player_ratings.rd) is downstream of this.

USTA NTRP — the fair-league classifier (USA, 1978)

  • Origin/philosophy: built by USTA/USPTA to classify players "for more compatible matches" — the league-play machine for the world's largest recreational tennis market. NTRP does not want to measure you precisely; it wants your Tuesday-night 3.5 league to be fun and fair.
  • Construction: a public coarse scale (1.5–7.0 in 0.5 bands, anchored by written "General Characteristics") sitting on a hidden two-decimal dynamic rating recalculated after every match but never published. One number per player across all formats. Entry via self-rate questionnaire; integrity via a human governance layer — auto-granted up-appeals, committee down-appeals, and the three-strike auto-disqualification that polices sandbagging (a named cultural problem).
  • Blind spots: coarse bands frustrate band-edge players; one-number design can't speak to singles/doubles gaps; sectional pools drift apart (a 4.0 in one region ≠ another). The strike system exists because any system whose bands gate entry creates an incentive to stay low — the deepest lesson NTRP offers us.

FFT classement — the aspiration pyramid (France, ~1920s)

  • Origin/philosophy: among the oldest continuously running rating traditions in sport, built for tournament seeding in a country where the classement is part of tennis identity. It is an ordinal ladder (NC → 40 → 30 → 15 → 5/6 … 0 → séries), climbed by beating people above you. Losses gate bonuses but never subtract — the ratchet reflects the philosophy: the classement is a career, not a snapshot.
  • Construction: points per win scaled by the opponent's échelon gap, minimum win counts per échelon, monthly rolling recompute since Oct 2022 (a genuine epoch break from the historical annual reset).
  • Blind spot/strength: non-numeric, direction-confusing to outsiders, and slow downward — but its competitiveness is legendary: UTR's founder explicitly cited FFT-style level-based events as producing far more competitive matches than US junior tennis, and built UTR to generalize that.

UTR — the market disruptor (private, 2008)

  • Origin/philosophy: Dave Howell, a Virginia pro, wanted one global, up-to-date scale where any two players anywhere could be compared — directly inspired by FFT's classement but rebuilt as a Family-A estimator. Its killer app was college recruiting (coaches comparing a Serbian junior to a Californian one), and its product creed is "level-based play": competitive matches are more fun, so measure precisely and match tightly.
  • Construction: 1.00–16.50 at two decimals; rolling 30-match/12-month window; per-match rating driven by opponent gap AND games-won margin (score margin matters — a philosophical break from win/loss purism); separate singles/doubles numbers; verified vs unverified tiers; daily recompute.
  • Blind spots: recency-heavy (an injury layoff moves you); margin-sensitivity creates run-up-the-score incentives; privately owned — the scale's governance is a company's product decisions.

ITF WTN — the federation's answer (global, 2021)

  • Origin/philosophy: the ITF plus national federations, mission stated plainly at launch: break down "one of sport's key barriers to participation — uneven match-ups." WTN is a participation-growth instrument wearing a rating's clothes: get every player a number so every player can find a fair match, worldwide, under federation governance (a strategic response to UTR's private land-grab).
  • Construction: 40 (beginner) → 1 (elite) — deliberately reversed so beginners "count down" as they improve; separate singles and doubles numbers (explicitly not comparable); Glicko-2 engine scoring at set level; a public confidence percentage (RD in a friendly costume) with a 70% "verified" threshold; free via federations; public GraphQL API. Adopted outright by the LTA (replacing its old ratings), the ITA (US college), JTA (Japan, in progress).
  • Blind spots: algorithm revisions reposition players in place (the scale itself isn't stable); coverage skews to federation-registered players — thin in exactly the casual populations apps like ours serve.

DTB Leistungsklasse — the club-league utility (Germany, ~2007)

  • Origin/philosophy: grew out of regional systems (Saarland/Rheinland-Pfalz) and was nationalized to solve a specific, very German problem: Medenspiele team lineups must be ordered by strength, so strength must be a single administrable number. The LK is infrastructure, not identity — and it deliberately incentivizes activity.
  • Construction: 25 (beginner) → 1 (best); one number across ALL formats (singles, doubles, league, tournaments feed the same LK); per-win improvement = points × age factor / hurdle, where the hurdle grows steeply toward LK 1 (physics-like resistance at the top); and the famous monthly decay cron — every inactive player drifts +0.1 toward 25 on the last Wednesday of each month. Play or sink.
  • Blind spots: no discipline split at all (your doubles-carried LK follows you into singles brackets); no confidence concept — the number is simply administratively true.

KR ladders (KTA/KATO/KASTA/테니스타운) — credential gates (Korea)

  • Philosophy: documented in full in our own research — these are licensing tiers with expiry, built for bracket fairness and anti-sandbagging in a tournament culture with formal elite (선출) segregation. Standing = qualification history; promotion mechanics are one-way doors (개나리 graduation, 챌린저 exclusion); "ratings" in the statistical sense don't exist — the KTA pair-handicap point tables (10.0→1.0) are the closest thing, and they're licensing math, not inference.
  • Why it matters to us: this is the vocabulary our KR users actually speak, which is why our credential rows cite it — but it is categorically Family B, and mapping it to any Family-A number is calibration across philosophies, only defensible as labeled bands.

Japan — the fragmented market (JTA/JOP + 草トー + WTN incoming)

No unified amateur system: JOP ranking serves the competitive registered class; the huge casual 草トーナメント scene runs on per-organizer labels (オープン/A/B/C級, platform-specific ladders like テニスベア's Lv.1–9) that are explicitly non-standardized; JTA is adopting WTN top-down. Philosophy vacuum at the amateur middle — which is an opportunity for an app with its own credible internal rating.

Playtomic — the product-led rating (padel, global)

0.0–7.0 Elo-style level with a visible reliability score, seeded by questionnaire, tuned per country. Philosophy: the rating exists to make the product work (auto-matched competitive games), not to be an institution. The closest structural cousin to what a club app's internal rating actually is — and evidence that a self-owned rating can become a de facto standard if the product wins.

How ours compares

Our architecture, named honestly, is: a Family-A engine wearing a Family-B display, in a Family-B market.

AxisOursClosest kinNotes
EngineELO with rd per trackWTN (Glicko-2)Same lineage; we keep uncertainty like WTN, unlike LK/FFT
Public shape5 coarse tiers (브론즈→그랜드마스터) over a hidden precise ratingNTRP exactlyNTRP proved this structure for 47 years: precise engine, coarse honest display, less band-anxiety
Discipline topologythree fully independent tracks (단식/복식/혼복)WTN/UTR (split S/D) — we go furtherThe most purist position in the market; nobody else runs an independent mixed track. US mixed pair-sum stays an eligibility rule, which our interop model handles as a declaration
Vocabulary layerPer-market credential/scale labels (KR divisions today)UTR-app side-by-side patternTiers never are NTRP/divisions — they're annotated with them
GovernanceAuto only (no strike/appeal layer yet)NTRP's committee layer is the warningThe moment tiers gate session entry (they already can), sandbagging incentives exist — NTRP's strike system is the precedent to study when we see it
DecayNone yet (rd grows implicitly stale)LK's cron vs Glicko's RD-inflationA deliberate open decision: we currently neither decay ratings (LK-style) nor formally inflate rd on inactivity (Glicko-style). Worth deciding before leaderboards matter
Seedingseed_rating + self-declared credentialsPlaytomic questionnaire; NTRP self-rateSame pattern as everyone; our verified-credential path (elite RPC, external ratings) is the NTRP "computer rating" analog

The three positions we hold that the research validates:

  1. Hidden-precise + public-coarse (NTRP's proven structure) — band-edge honesty without false precision.
  2. Uncertainty as a first-class citizen (Glicko/WTN lineage) — rd already in the schema.
  3. Independent tracks — stricter than anyone; defensible because consolidation is exactly where NTRP's one-number and LK's all-formats designs generate their known complaints.

The three warnings the world's systems give us:

  1. Sandbagging follows gating (NTRP): when tier bands gate entry, some players want low ratings. Our reliability gates + host authority are the current defense; a formal review/strike mechanism is the eventual one.
  2. Decay is a decision, not a default (LK vs Glicko): choose activity-incentive decay, uncertainty inflation, or neither — explicitly.
  3. Scales you anchor to will move under you (WTN revisions, FFT 2022): which is why the interop model versions every crosswalk and stamps every observation.

What "side by side" means concretely

A profile shows: our tier (earned, internal, the only thing we compute) · our band expressed in the market's primary scale (labeled inference — "복식 NTRP 3.5~4.0 수준" / someday "WTN ~25 Zone") · the user's verified external ratings (observations with provenance and dates — "UTR 5.2 (verified, 2026-05)") · credentials (Family-B facts — "오픈부 입상"). Four different epistemic objects, four different visual treatments, never merged into one number. That's the whole philosophy in one sentence: we compute one thing, translate honestly, display the rest as what it is.

Markdown remains the source of truth. Run yarn docs:check before handoff.