⚽ ConvocadosDocsllms.txtGitHub

The cost of certainty: why Elo can't tell a rookie from a veteran

Two players rated 1500 are not the same player. One number hides that.

The previous page treated every rating as an exact measurement. It isn't. A rating is an estimate, and an estimate should carry an error bar. Elo has none — so it moves a brand-new player and a ten-year veteran as if both were equally known. That is the second structural flaw, and it hits Convocados the moment a Season starts with fresh members.

1 · The missing error bar

Glicko's fix (Glickman, 1995) is to give each player a second number: therating deviation (RD), the system's uncertainty about their rating. A rating is then reported as an interval — a 95% confidence range ofr ± 2·RD.

Same rating, different confidence
Games played · 0
Current rating · 1500
8002200

Rating deviation RD = 350 (95% confidence: 800 – 2200)

Next win moves you: +175.0 pts (beating an equal 1500)

Starting RD is 350 (wide). Consistent results shrink it; each 95% range should feel visibly narrower. RD grows again during inactivity, because a rating frozen for a year is stale evidence.

The update itself respects RD: the bigger your uncertainty, the further one game can move you. This is why a provisional player belongs at the top of a ladder by the end of a night — the system is learning fast because it knows little.

2 · Uncertainty that breathes

Two refinements matter for a weekly hobby league:

Where Convocados stands. Season Rank substitutes a crude uncertainty rule — K=64 for a player's first three counted games, K=32after — and then freezes the payout forever. A real RD would make the transition smooth, widen after a layoff, and let the provisional window be a function of evidence rather than a hard-coded count.

3 · Many people, one team

Elo and Glicko are head-to-head. Convocados is 5-a-side: one result, ten players, uneven contribution. Two families handle this:

Both beat Elo at team settings, but neither knows who was on the bench or how minutes were shared — a caveat Convocados must handle in product, not maths. Team rating is a task for the balancing layer, not necessarily for the motivational ladder.

Check yourself

Rating deviation measures, for a rating, the system's…

Confidence in the estimate. RD is the system's uncertainty, reported as a confidence interval.

A player who returns after a long layoff should have their rating deviation…

Widened while inactive. Skill may have drifted while away, so the old evidence is stale.

TrueSkill's main advantage over Elo in a team sport is…

Learning from team outcomes. Team results let the model infer each member's contribution.

Primary sources

Work throughGlickman's worked Glicko-2 example— it is short and shows every parameter. For teams, skim theTrueSkill paper, then read theOpenSkill paper for what you'd actually ship.