Soft Condorcet Optimization for Ranking of General Agents (Extended Abstract)

Soft Condorcet Optimization for Ranking of General Agents (Extended Abstract)

Marc Lanctot, Kate Larson, Michael Kaisers, Quentin Berthet, Ian Gemp, Manfred Diaz, Roberto-Rafael Maura-Rivero, Yoram Bachrach, Anna Koop, Doina Precup

Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
Sister Conferences Best Papers. Pages 8272-8277. https://doi.org/10.24963/ijcai.2026/925

Driving progress of AI models and agents requires comparing their performance on standardized benchmarks; for general agents, individual performances must be aggregated across a potentially wide variety of different tasks. In this extended abstract, we describe a ranking scheme inspired by social choice frameworks, called Soft Condorcet Optimization (SCO), to compute the optimal ranking of agents: the one that makes the fewest mistakes in predicting the agent comparisons in the evaluation data. This optimal ranking is the maximum likelihood estimate when evaluation data (which we view as votes) are interpreted as noisy samples from a ground truth ranking, a solution to Condorcet's original voting system criteria. SCO ratings are maximal for Condorcet winners when they exist, which we show is not necessarily true for the classical rating system Elo. In practice, SCO serves as an accurate approximation to the Kemeny-Young voting method, excels in the sparse data regime, and provides the best approximation to the optimal ranking compared to every baseline on a Diplomacy player ranking problem with 31,094 games and 52,958 agents.
Keywords:
Machine Learning: Evaluation
AI: Agent-based and Multi-agent Systems
Machine Learning: Optimization
Game Theory and Economic Paradigms: Computational social choice