In preparation for meeting with a national sport association towards a Bayesian approach to ranking (I can only confirm this not for volleyball!), I was searching for advanced studies of the Elo (not ELO!) rating system and came across this 2024 arXival by Sam Olesker-Taylor (U Warwick) and Luca Zanetti. Where they analyse the online behaviour of the player ratings and their convergence (in Wasserstein distance), with fairly intricate proofs.
“Elo is not a reversible Markov chain and, while it has a unique stationary distribution, assuming a minor and natural condition, it does not converge to it in total variation.”
The analysis is based on the Bradley–Terry–Luce model of the probability of player i winning a game against player j. Which amounts to a logit or sigmoïd transform of the difference between the players’ true ratings. The practical Elo updates amounts to one stochastic gradient step for the associated likelihood. The authors also propose a random allocation of players into pairs that achieves optimal convergence to the true rating, by maximising the spectral gap of the allocation matrix. Although I did not spot how to derive it in practice.