Ranking Reasoning LLMs under Test-Time Scaling

Mohsen HaririMichael HinczewskiJing MaVipin Chaudhary

Prefer the short version? Read the overview of this paper.

Abstract

Test-time scaling evaluates reasoning LLMs by sampling multiple outputs per prompt, but ranking models in this regime remains underexplored. We formalize dense benchmark ranking under test-time scaling and introduce Scorio, a library that implements statistical ranking methods such as paired-comparison models, item response theory (IRT) models, voting rules, and graph- and spectral-based methods. Across 2020 reasoning models on four Olympiad-style math benchmarks (AIME'24, AIME'25, HMMT'25, and BrUMO'25; up to N=80N=80 trials), most full-trial rankings agree closely with the Bayesian gold standard BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 (mean Kendall's τb=0.93\tau_b = 0.93 - 0.950.95), and 1919 - 3434 methods recover exactly the same ordering. In the single-trial regime, the best methods reach τb≈0.86\tau_b \approx 0.86. Using greedy decoding as an empirical prior (BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N) reduces variance at N=1N=1 by 1616 - 52%52\%, but can bias rankings when greedy and stochastic sampling disagree. These results identify reliable ranking methods for both high- and low-budget test-time scaling. We release Scorio as an open-source library at GitHub (https://github.com/mohsenhariri/scorio. See Appendix J for API documentation.).

Introduction

Agreement between each method's full-trial ranking and the gold standard. Kendall's _b is computed between each method's ranking (at N=80 trials) and Bayes_ U @80 on an easier benchmark (BrUMO'25 , left) and the hardest benchmark (HMMT'25 , right). On BrUMO'25 , multiple methods achieve near-perfect or perfect agreement: Bayes_ R_0 @N and HodgeRank reach _b = 1.0, while Rasch MML achieves 0.997. On HMMT'25 , Bradley - Terry and HodgeRank maintain perfect agreement (_b = 1.0), but Bayes_ R_0 @N drops to 0.989 and 2 falls to 0.937. This divergence is consistent with the lower greedy - sampling alignment observed on harder benchmarks ().
Figure 1. Agreement between each method's full-trial ranking and the gold standard. Kendall's τb\tau_b is computed between each method's ranking (at N=80N=80 trials) and BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 on an easier benchmark (BrUMO'25, left) and the hardest benchmark (HMMT'25, right). On BrUMO'25, multiple methods achieve near-perfect or perfect agreement: BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N and HodgeRank reach τb=1.0\tau_b = 1.0, while Rasch MML achieves 0.9970.997. On HMMT'25, Bradley - Terry and HodgeRank maintain perfect agreement (τb=1.0\tau_b = 1.0), but BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N drops to 0.9890.989 and Pass@22 falls to 0.9370.937. This divergence is consistent with the lower greedy - sampling alignment observed on harder benchmarks (section 3.4).

Large language models (LLMs) are increasingly used as general-purpose reasoning systems for tasks such as programming and mathematical problem solving [1, 2]. Reliable evaluation is therefore essential. In many settings, what matters is not only an absolute score but also a ranking that supports model selection, deployment, and scientific comparison. This need is amplified by test-time scaling, which allocates additional inference compute by sampling multiple outputs per prompt and aggregating them, turning evaluation into a repeated-sampling problem [2, 3, 4].

Statistical ranking methods underpin two common LLM workflows. First, preference-based learning and alignment pipelines rely on human or model preferences over alternative responses, where the primitive observations are paired comparisons and downstream optimization depends on how those preferences are modeled and aggregated [5, 6]. Second, model comparisons are often communicated through leaderboards. Crowdsourced paired-comparison platforms such as Chatbot Arena collect head-to-head judgments and fit rating or paired-comparison models to produce public rankings [7], while benchmark-style evaluations rank models by task performance metrics such as Pass@kk [1]. Recent work has revisited the statistical foundations of LLM ranking in both preference-based settings [8] and benchmark settings, including IRT-style benchmarking [9]. Different ranking methods can produce noticeably different model orderings, and their agreement can vary with benchmark difficulty (figure 1).

A key practical distinction between these regimes is the representation of the data used for ranking. Preference-based evaluation typically yields a sparse and evolving comparison graph because only a subset of model pairs are compared and the model pool changes over time [7]. In contrast, benchmark evaluations produce dense outcomes for every model - question pair. For a fixed set of LL models and MM questions, we observe an outcome for every pair. Under test-time scaling, each model - question pair is evaluated with NN independent trials, producing a response tensor R∈{0,1}L×M×N\mathbf{R}\in\{0,1\}^{L\times M\times N}. This dense repeated-trial setting raises new methodological questions: Which ranking rule should be used when NN is small? How quickly do different ranking methods stabilize as NN grows? How do priors and uncertainty estimates affect ranking robustness?

In this work, we study performance-based ranking under test-time scaling. We formalize the dense benchmark setting through the response tensor R\mathbf{R}, evaluate ranking methods by their low-budget stability and convergence as the test-time budget increases, and implement the studied methods in Scorio.

We summarize our contributions as follows:

  • We formalize dense benchmark ranking under test-time scaling via R∈{0,1}L×M×N\mathbf{R}\in\{0,1\}^{L\times M\times N} and connect common ranking families through pointwise, pairwise, and setwise transformations of R\mathbf{R}.

  • We propose an evaluation protocol based on low-budget stability (agreement between rankings computed from subsampled trials and reference rankings) and convergence with increasing numbers of trials.

  • We compare a broad suite of ranking methods across 2020 reasoning models and four Olympiad-style math benchmarks (up to N=80N=80 trials), characterizing where method families agree and where they diverge.

  • We analyze Bayesian and uncertainty-aware ranking choices, including priors and conservative (quantile-based) scoring, and quantify their bias - variance trade-offs in low-trial regimes.

  • We release Scorio, a library implementing the ranking methods and Bayesian options.

Ranking Problem and Test-time Scaling

In classical statistical settings, there is no canonical theoretical ground truth or empirical gold standard against which competing ranking rules can be judged. Choosing among methods therefore usually requires additional modeling assumptions. Test-time scaling offers a useful alternative: because each model - question pair can be sampled repeatedly, it lets us evaluate ranking methods by how stable they are in low-budget settings and how quickly they converge as more trials are observed.

Statistical ranking methods are widely used in domains such as sports competitions (e.g., paired-comparison models and rating systems for head-to-head games) [10, 11, 12] and voting or collective decision-making [13, 14, 15]. In such settings, there are LL entities to be ranked (e.g., players, items, or models) over MM tasks (e.g., matches, questions, or instances). Test-time scaling adds a third dimension: NN, the number of i.i.d. samples generated for a fixed question m∈{1,…,M}m\in\{1,\dots,M\}. Repeated sampling lets us study two complementary properties. First, low-budget stability asks whether a ranking computed from a small number of trials agrees with a high-budget reference ranking. In our experiments, the low-budget case is N=1N=1: we subsample one trial per question, compute the ranking, repeat this over the available single-trial draws, and compare each ranking either with an empirical gold standard or with the same method's full-trial ranking. Second, convergence asks how quickly rankings computed from nn trials approach the full-trial ordering as nn increases from 11 to NN.

Gold Standard Rankings

Evaluation metrics widely used in test-time scaling, such as Pass@kk and Bayes@NN, can be analyzed through statistical properties such as bias. For instance, [1] derive an unbiased estimator for Pass@kk. As the number of trials NN grows, empirical estimates of these metrics concentrate around their population values, making metric-based rankings increasingly stable. In particular, for binary outcomes, BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N is order-equivalent to mean accuracy avg@NN [16], which motivates our use of the full-trial BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N ranking as an empirical accuracy-based gold standard.

This reasoning does not extend automatically to all ranking methods. Even as the number of questions MM or trials NN increases, different ranking methods need not converge to a unique limiting ordering, such as the one induced by average accuracy (Appendix C.1). Unlike evaluation metrics, ranking algorithms can emphasize different aspects of performance across tasks, players, or items. In section 3, we show that rankings induced by probabilistic models (e.g., Bradley - Terry) can differ from those induced by expected-performance metrics (e.g., mean accuracy or Bayesian estimates).

Given the absence of a universal gold standard for ranking methods, we use two target rankings for comparison. First, we define an empirical gold standard based on average performance over all trials with a large sample size (e.g., N=80N=80). This target captures aggregate performance across tasks and trials while allowing ties. This choice is justified for several reasons: (a) the ranking induced by average performance is order-equivalent to the ranking induced by Bayesian estimation with a uniform prior (BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N); (b) when NN is large, average performance is among the most stable ranking rules relative to the alternatives (section 3.1); and (c) it is easy to interpret, widely used in practice, and yields absolute performance values.

The second target ranking is the ordering produced by a method itself (method@8080) when all available trials are aggregated. This target lets us assess a method's self-consistency and convergence as more data become available.

Representation

We consider LL models evaluated on a benchmark of MM questions under test-time scaling, generating NN i.i.d. trials per model - question pair. Let L={1,…,L}\mathcal{L}=\{1,\dots,L\} index models and Q={1,…,M}\mathcal{Q}=\{1,\dots,M\} questions; for each question we observe NN independent trials indexed by n∈{1,…,N}n\in\{1,\dots,N\}. For each (l,m,n)∈L×Q×{1,…,N}(l,m,n)\in \mathcal{L}\times\mathcal{Q}\times\{1,\dots,N\} we observe a binary outcome

Rlmn∈{0,1},R_{lmn} \in \{0,1\}, (1)

where Rlmn=1R_{lmn}=1 if model ll solves question mm on trial nn. We collect these outcomes in a response tensor R∈{0,1}L×M×N\mathbf{R}\in\{0,1\}^{L\times M\times N}. When N=1N=1, this reduces to the standard single-run benchmark setting. Unlike crowdsourced paired-comparison datasets (e.g., Chatbot Arena [7]), where the primitive observations are model - model outcomes on a possibly sparse comparison graph, our benchmark setting produces outcomes for every model - question pair. We discuss the Arena sparse-comparison counterpart in appendix G. We therefore take R\mathbf{R} as the primitive object; all ranking methods we study use R\mathbf{R} as input, but they differ in the representations on which they operate after transforming or aggregating it.

Pointwise (model - question) representation.

Define the per-question solve rate

p^lm:=1N∑n=1NRlmn,\widehat{p}_{lm} := \frac{1}{N}\sum_{n=1}^N R_{lmn}, (2)

and the overall mean accuracy p^l:=1M∑m=1Mp^lm\widehat{p}_{l} := \frac{1}{M}\sum_{m=1}^M \widehat{p}_{lm}. Pointwise and IRT-style methods operate on the matrix P^=[p^lm]∈[0,1]L×M\widehat{\mathbf{P}}=[\widehat{p}_{lm}]\in[0,1]^{L\times M} (or on its row means), optionally reweighting questions (e.g., inverse-difficulty weighting [17]). Classical IRT models infer latent abilities from this representation [18, 19], and have recently been applied to LLM benchmarking [9]. When N>1N>1, the trial axis corresponds to repeated Bernoulli observations; likelihood-based models (including IRT) can equivalently work with the sufficient statistic klm:=∑nRlmnk_{lm}:=\sum_{n}R_{lmn}, yielding a binomial-response formulation [20, 21]. Related repeated-measures and longitudinal IRT extensions are also well studied [22, 23]. Evaluation-metric rankings (e.g., Pass@kk and Bayes@NN) additionally use the per-question trial multiset {Rlm1,…,RlmN}\{R_{lm1},\dots,R_{lmN}\} (equivalently the count ∑nRlmn\sum_n R_{lmn}) to compute per-question metrics before aggregating across mm [1].

Pairwise (win/tie) representation.

Many classical ranking methods reduce R\mathbf{R} to pairwise outcomes. For a pair of models (i,j)∈L2(i,j)\in\mathcal{L}^2 we define win and tie counts

Wij:=∑m=1M∑n=1N1{Rimn=1,  Rjmn=0},\begin{aligned}W_{ij} &:= \sum_{m=1}^M\sum_{n=1}^N \mathbf{1}\{R_{imn}=1,\; R_{jmn}=0\},\end{aligned} (3)
Tij:=∑m=1M∑n=1N1{Rimn=Rjmn},\begin{aligned}T_{ij} &:= \sum_{m=1}^M\sum_{n=1}^N \mathbf{1}\{R_{imn}=R_{jmn}\},\end{aligned} (4)

so that, in our fully observed setting, Wij+Wji+Tij=MNW_{ij}+W_{ji}+T_{ij}=MN for all i≠ji\neq j. Equivalently, we can form an undirected comparison graph G=(V,E)G=(V,E) with vertex set V=LV=\mathcal{L} and edge set E={{i,j}:Wij+Wji+Tij>0}E=\{\{i,j\}:W_{ij}+W_{ji}+T_{ij}>0\}, and store (Wij,Wji,Tij)(W_{ij},W_{ji},T_{ij}) on each edge. In our benchmark setting EE is the complete graph (every pair is compared MNMN times), whereas in interactive evaluation settings EE is typically sparse and one assumes GG is connected. The matrices W=[Wij]\mathbf{W}=[W_{ij}] and T=[Tij]\mathbf{T}=[T_{ij}] define a weighted comparison graph over models. Probabilistic paired-comparison models (e.g., Bradley - Terry and tie extensions [10, 24, 25]) and voting rules (e.g., Borda and Copeland [13, 26]) use these aggregated counts; graph- and spectral-based methods (e.g., PageRank, Rank Centrality, HodgeRank, SerialRank, AlphaRank, and Nash-based ranking [27, 28, 29, 30, 31, 32]) further transform (W,T)(\mathbf{W},\mathbf{T}) into Markov chains or skew-symmetric edge flows, typically via edge weights based on empirical win rates such as P^i≻j=(Wij+12Tij)/(Wij+Wji+Tij)\widehat{P}_{i\succ j}=(W_{ij}+\tfrac{1}{2}T_{ij})/(W_{ij}+W_{ji}+T_{ij}). Sequential rating systems (e.g., Elo and TrueSkill [11, 33]) instead process the underlying stream of pairwise "matches" induced by each question - trial (m,n)(m,n).

Listwise or setwise representation.

For each question - trial (m,n)(m,n) we define the winning set Umn:={l∈L:Rlmn=1}U_{mn}:=\{l\in\mathcal{L} : R_{lmn}=1\} and the losing set L∖Umn\mathcal{L}\setminus U_{mn}, which induces a two-level partial order: all winners tie above all losers. Setwise or listwise models (e.g., Plackett - Luce [34, 35] and Davidson - Luce [36]) operate directly on the collection of events {(Umn,L∖Umn)}m,n\{(U_{mn},\mathcal{L}\setminus U_{mn})\}_{m,n}, discarding degenerate events with Umn=∅U_{mn}=\emptyset or Umn=LU_{mn}=\mathcal{L}. In our binary two-level setting, Plackett - Luce likelihoods collapse to functions of pairwise win counts (cf. the MM formulation for generalized Bradley - Terry and Plackett - Luce likelihoods [37]), whereas Davidson - Luce explicitly models within-set ties.

Bayesian Approaches in Ranking

Many ranking methods can be viewed as probabilistic models with latent parameters θ\theta (e.g., model strength and, optionally, question difficulty). Given observations R\mathbf{R} (or derived representations such as pairwise counts; section 2.2), inference reduces to estimating θ\theta from a likelihood p(R∣θ)p(\mathbf{R}\mid\theta). We consider maximum likelihood estimation (MLE), maximum a posteriori (MAP), and expected a posteriori (EAP), and discuss how uncertainty can be propagated to rankings [38]. Although MLE is not Bayesian, we include it as a standard baseline for likelihood-based ranking models.

Maximum likelihood estimation (MLE).

The maximum likelihood estimate is

θ^MLE∈arg⁡max⁡θ  p(R∣θ),\hat{\theta}_{\mathrm{MLE}} \in \arg\max_{\theta} \; p(\mathbf{R} \mid \theta), (5)

which yields a point estimate without requiring a prior. MLE is attractive for its simplicity, but in paired-comparison and IRT-like models it can be unstable under (near-)separation or weak identification, which motivates priors in MAP and EAP.

Maximum a posteriori (MAP).

MAP incorporates prior information p(θ)p(\theta) and estimates the posterior mode:

θ^MAP∈arg⁡max⁡θ  p(R∣θ) p(θ).\hat{\theta}_{\mathrm{MAP}} \in \arg\max_{\theta} \; p(\mathbf{R} \mid \theta)\, p(\theta). (6)

Equivalently, MAP is a penalized MLE in which −log⁡p(θ)-\log p(\theta) acts as a regularizer; priors can improve stability in paired-comparison and IRT-style models [39, 40]. We can also construct empirical priors from auxiliary evaluation runs. For example, a prior outcome tensor R0\mathbf{R}_0 (e.g., one greedy decode per question) can be used to regularize stochastic trials (EmpiricalPrior in Scorio) [16].

Expected a posteriori (EAP).

EAP uses the posterior mean as the point estimate:

θ^EAP:=E[θ∣R],\hat{\theta}_{\mathrm{EAP}} := \mathbb{E}[\theta \mid \mathbf{R}], (7)

which is Bayes-optimal under squared-error loss [38]. Compared with MAP, EAP accounts for posterior mass beyond the mode and typically requires approximation or sampling. EAP is common in latent-trait settings such as IRT and adaptive testing [41].

Interval estimates and conservative ranking.

Bayesian methods naturally yield credible intervals (posterior quantiles) for each θl\theta_l, while frequentist analyses can produce approximate confidence intervals for θ^MLE\hat{\theta}_{\mathrm{MLE}} via bootstrap resampling of questions or trials. Interval estimates are especially useful because ranking is sensitive to near ties: rather than ranking by point estimates alone, one can rank conservatively using a lower credible or confidence bound (LCB), or report pairwise superiority probabilities Pr⁡(θi>θj∣R)\Pr(\theta_i > \theta_j \mid \mathbf{R}). Metric-level Bayesian estimators such as Bayes@NN provide both a posterior mean and uncertainty, enabling rankings by posterior mean or by a chosen posterior quantile. Bayes@NN also supports incorporating prior outcomes R0\mathbf{R}_0 (e.g., one greedy decode per question) as pseudo-counts in the posterior, which is complementary to using R0\mathbf{R}_0 to define empirical priors for MAP in parametric ranking models. Our implementation in Scorio supports both credible-interval ranking via Bayes@NN and empirical priors via EmpiricalPrior for MAP estimation.

Experiments

We evaluate 7272 ranking methods (Appendix J.2) on four Olympiad-style math benchmarks: AIME'24, AIME'25, HMMT'25, and BrUMO'25, each with M=30M=30 questions. We use L=20L=20 reasoning LLMs (full list in table 22). For each model - question pair, we collect N=80N=80 independent trials via top-pp sampling, yielding a response tensor R∈{0,1}20×30×80\mathbf{R}\in\{0,1\}^{20\times 30\times 80}. We also collect a single greedy-decoding output per question (R0\mathbf{R}_0) to serve as an empirical prior. Detailed generation, sampling, and reproducibility settings appear in appendix I; the library API is documented in appendix J.

Gold Standard Ranking

Following section 2.1, we define the gold-standard ranking as BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80, the Bayesian posterior-mean estimator with a uniform prior computed from all N=80N=80 trials. This choice is order-equivalent to avg@8080 (mean correctness over all MM questions and all N=80N=80 trials, with ties allowed) and yields an interpretable accuracy-based target. Empirically, when each of our 7272 ranking methods is computed using all 8080 trials, the resulting orderings agree closely with BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 (table 1): across benchmarks, the average Kendall's τb\tau_b between BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 and the other methods is 0.930.93 - 0.950.95 (median 0.950.95 - 0.990.99), and 1919 - 3434 methods recover exactly the same ordering (τb=1\tau_b=1). The largest deviations come from a small set of voting rules (e.g., minimax and Nanson variants) and difficulty-weighted baselines, with minimum τb\tau_b values of 0.680.68 - 0.790.79 depending on the benchmark. Although BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N is order-equivalent to avg@NN, we prefer the Bayesian formulation because it supports priors (e.g., BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N) and uncertainty estimates.

Table 1. Agreement between the gold-standard ranking (BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80) and each other ranking method, measured by Kendall's τb\tau_b, when all methods are computed from the full N=80N=80 trials. Statistics are computed over the other 7171 methods; "Combined" pools all benchmarks.
BenchmarkMeanMedianMin#(τb=1\tau_b=1)#(τb≥0.95\tau_b\ge 0.95)
AIME'240.9410.9890.6822040
AIME'250.9340.9470.7711929
HMMT'250.9500.9890.7583444
BrUMO'250.9540.9680.7892649
Combined0.9620.9890.7482253

Ranking-Method Stability

To compare ranking methods in the low-budget regime, we set N=1N=1 by subsampling one of the 8080 trials per question and recomputing the rankings. For each method, we report Kendall's τb\tau_b averaged over the 8080 single-trial draws (mean ±\pm std). Since the Pass@kk family requires at least two trials to differ from mean accuracy, the N=1N=1 comparisons below cover the remaining 6969 methods.

Gold-standard agreement.

We first rank methods by agreement with the empirical gold standard (BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80). Across AIME'24, AIME'25, and BrUMO'25, BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N performs best, achieving τb=0.779±0.034\tau_b=0.779\pm0.034, 0.798±0.0450.798\pm0.045, and 0.858±0.0280.858\pm0.028, respectively (table 2). On HMMT'25, the hardest benchmark (see appendix B), the greedy prior no longer helps, and the best score is shared by a 2121-method equivalence class (BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N and several graph- and voting-based methods), with τb=0.790±0.053\tau_b=0.790\pm0.053. When all benchmarks are pooled (Combined), the same 2121-method class attains τb=0.865±0.049\tau_b=0.865\pm0.049, while BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N drops to τb=0.786±0.031\tau_b=0.786\pm0.031 (table 17).

Self-consistency and convergence.

Next, we evaluate each method against its own full-trial ranking (method@8080), which summarizes convergence from N=1N=1 to N=80N=80. Rasch MML with LCB scoring is the most self-consistent on AIME'24, AIME'25, and HMMT'25, with τb=0.804±0.051\tau_b=0.804\pm0.051, 0.834±0.0540.834\pm0.054, and 0.810±0.0560.810\pm0.056 (table 2); BrUMO'25 again favors BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N (0.858±0.0280.858\pm0.028). On the Combined benchmark, the most self-consistent method is Nanson's rule with tie averaging (0.892±0.0500.892\pm0.050), followed by Rasch MML (LCB) (0.883±0.0370.883\pm0.037), whereas several minimax variants are among the least self-consistent (down to 0.765±0.0450.765\pm0.045; table 18). High self-consistency does not imply strong agreement with the gold standard: Nanson (avg ties) ranks first in self-consistency on Combined but has substantially lower gold-standard agreement (0.807±0.0360.807\pm0.036; table 17).

Table 2. Best-performing ranking methods in the low-budget regime (N=1N=1) under two targets: (i) agreement with the gold standard (BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80) and (ii) self-consistency with each method's own full-trial ranking (method@8080). Kendall's τb\tau_b is averaged over 8080 single-trial draws; †\dagger denotes a 21-way tie for best gold-standard agreement (see table 17). Pass@kk variants are excluded at N=1N=1 because they require N≥2N \ge 2. Method identifiers correspond to the APIs listed in appendix J.2.
BenchmarkBest vs. gold standardτb\tau_bBest self-consistency (vs. method@80)τb\tau_b
AIME'24BayesR0@1\mathrm{Bayes}_{\mathbf{R}_0}@10.779±0.0340.779 \pm 0.034Rasch MML LCB (rasch_mml_credible)0.804±0.0510.804 \pm 0.051
AIME'25BayesR0@1\mathrm{Bayes}_{\mathbf{R}_0}@10.798±0.0450.798 \pm 0.045Rasch MML LCB (rasch_mml_credible)0.834±0.0540.834 \pm 0.054
HMMT'25Bayes@11 †\dagger0.790±0.0530.790 \pm 0.053Rasch MML LCB (rasch_mml_credible)0.810±0.0560.810 \pm 0.056
BrUMO'25BayesR0@1\mathrm{Bayes}_{\mathbf{R}_0}@10.858±0.0280.858 \pm 0.028BayesR0@1\mathrm{Bayes}_{\mathbf{R}_0}@10.858±0.0280.858 \pm 0.028
CombinedBayes@11 †\dagger0.865±0.0490.865 \pm 0.049Nanson avg ties (nanson_rank_ties_average)0.892±0.0500.892 \pm 0.050

Bootstrapped Model-Pool Robustness

The preceding N=1N=1 results use the full set of 2020 models. To test whether those conclusions depend on the evaluation pool, we repeat the low-budget analysis on bootstrapped model pools of size 55, 1010, and 1515. For each bootstrap subset, we recompute the full-trial rankings, use the subset-specific avg@8080 ordering as the gold-standard target, and compare each method's 8080 single-trial rankings against two references: (i) the subset-specific avg@8080 ordering and (ii) its own subset-specific full-trial ranking (method@8080). We aggregate 10001000 bootstrap subsets for each benchmark - size setting.

Easy and medium benchmarks preserve the original winner.

On AIME'24, AIME'25, and BrUMO'25, BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N remains the best representative method under both targets at all three model-pool sizes (table 3). The mean score changes only slightly with pool size: on AIME'24, gold-standard agreement moves from 0.7690.769 to 0.7800.780 and self-consistency from 0.7730.773 to 0.7850.785 as the pool size increases from 55 to 1515 models; on AIME'25, the corresponding ranges are 0.7970.797 - 0.8020.802 and 0.8030.803 - 0.8090.809; on BrUMO'25, BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N stays near 0.8540.854 - 0.8580.858 for both targets. On BrUMO'25, this advantage also becomes more decisive as the pool grows: the fraction of subsets where BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N is the top-scoring method rises from about 0.690.69 at k=5k=5 to 0.980.98 - 0.990.99 at k=15k=15.

Harder benchmarks remain tie-rich.

The harder settings behave differently. On HMMT'25 and on the Combined benchmark, the top score is not unique: for agreement with avg@8080, an equivalence class of 2929 - 3030 methods shares the best mean, while for method@8080 the tied class still contains 1313 - 1414 methods. We report avg (avg@NN, order-equivalent to BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N) as a representative member of these tied classes. The tied optimum is essentially flat across pool size, staying near 0.7880.788 - 0.7900.790 on HMMT'25 and 0.8630.863 - 0.8660.866 on Combined. This mirrors the full-model analysis in section 3.2: once the benchmark is difficult or pooled across heterogeneous tasks, many pointwise, voting, and graph-based methods become empirically indistinguishable.

Larger pools mainly reduce between-subset variance.

The primary effect of increasing the model-pool size is to reduce dispersion across subsets rather than to shift the mean systematically (table 3). For the best method under the avg@8080 target, the across-subset standard deviation falls from 0.2090.209 to 0.0570.057 on AIME'24, from 0.1440.144 to 0.0380.038 on AIME'25, from 0.1140.114 to 0.0330.033 on HMMT'25, from 0.1360.136 to 0.0320.032 on BrUMO'25, and from 0.0840.084 to 0.0230.023 on Combined when moving from 55 to 1515 models. Thus, the qualitative recommendation is stable under moderate changes to the model pool: larger pools mainly make the same conclusion more certain.

Table 3. Bootstrapped model-pool results in the low-budget regime (N=1N=1). For each model-pool subset, we compute Kendall's τb\tau_b over the 8080 single-trial rankings against two targets: the subset-specific gold standard avg@80 and each method's own subset-specific full-trial ranking (method@80). The table reports the mean and standard deviation of the subset-level mean score across bootstrap model pools for each subset size.
BenchmarkPoolBest Methodτb\tau_b vs Target
avg@80method@80
AIME'245BayesR0@1\mathrm{Bayes}_{\mathbf{R}_0}@10.769±0.2090.769 \pm {\scriptstyle 0.209}0.773±0.2070.773 \pm {\scriptstyle 0.207}
100.776±0.1070.776 \pm {\scriptstyle 0.107}0.781±0.1050.781 \pm {\scriptstyle 0.105}
150.780±0.0570.780 \pm {\scriptstyle 0.057}0.785±0.0570.785 \pm {\scriptstyle 0.057}
AIME'255BayesR0@1\mathrm{Bayes}_{\mathbf{R}_0}@10.802±0.1440.802 \pm {\scriptstyle 0.144}0.809±0.1440.809 \pm {\scriptstyle 0.144}
100.797±0.0710.797 \pm {\scriptstyle 0.071}0.803±0.0730.803 \pm {\scriptstyle 0.073}
150.798±0.0380.798 \pm {\scriptstyle 0.038}0.804±0.0400.804 \pm {\scriptstyle 0.040}
HMMT'255Bayes@110.788±0.1140.788 \pm {\scriptstyle 0.114}0.788±0.1140.788 \pm {\scriptstyle 0.114}
100.789±0.0590.789 \pm {\scriptstyle 0.059}0.789±0.0590.789 \pm {\scriptstyle 0.059}
150.790±0.0330.790 \pm {\scriptstyle 0.033}0.790±0.0330.790 \pm {\scriptstyle 0.033}
BrUMO'255BayesR0@1\mathrm{Bayes}_{\mathbf{R}_0}@10.854±0.1360.854 \pm {\scriptstyle 0.136}0.854±0.1360.854 \pm {\scriptstyle 0.136}
100.856±0.0620.856 \pm {\scriptstyle 0.062}0.856±0.0620.856 \pm {\scriptstyle 0.062}
150.858±0.0320.858 \pm {\scriptstyle 0.032}0.858±0.0320.858 \pm {\scriptstyle 0.032}
Combined5Bayes@110.863±0.0840.863 \pm {\scriptstyle 0.084}0.863±0.0840.863 \pm {\scriptstyle 0.084}
100.866±0.0420.866 \pm {\scriptstyle 0.042}0.866±0.0420.866 \pm {\scriptstyle 0.042}
150.864±0.0230.864 \pm {\scriptstyle 0.023}0.864±0.0230.864 \pm {\scriptstyle 0.023}

Effect of Empirical Priors

Empirical priors use auxiliary evaluation signals to stabilize low-budget rankings. In our setting, the signal is a single greedy decode, R0\mathbf{R}_0. We incorporate R0\mathbf{R}_0 into Bayes@NN, yielding BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N, and compare it with the uniform-prior variant BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N. We evaluate both variants by their agreement with the gold-standard ranking BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80. For each NN, we compute Kendall's τb\tau_b between the induced model ranking and BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 and report the mean and standard deviation over 5050 resampled datasets.

Gold-standard agreement of Bayes_ U @N (blue) and Bayes_ R_0 @N (red) as a function of N across benchmarks. Shaded regions show 1 standard deviation over 50 resampled datasets.
Figure 2. Gold-standard agreement of BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N (blue) and BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N (red) as a function of NN across benchmarks. Shaded regions show ±1\pm 1 standard deviation over 5050 resampled datasets.
Empirical priors reduce variance at low NN.

Across all benchmarks, BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N yields more stable low-NN rankings than BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N. At N=1N=1, the standard deviation of τb\tau_b decreases by 1616 - 52%52\% depending on the benchmark (table 4, figure 7). This advantage shrinks quickly as NN increases (figure 2), consistent with the prior contributing only O(1)O(1) pseudo-counts per question.

Table 4. Dataset difficulty (mean accuracy), greedy - sampling alignment (τG-S\tau_{\text{G-S}}), and the effect of the greedy empirical prior at N=1N=1. Δτ\Delta\tau is the difference in gold-standard agreement (greedy minus uniform), and Std. Red. is the relative reduction in the standard deviation of τb\tau_b.
BenchmarkDifficultyτG-S\tau_{\text{G-S}}Δτ\Delta\tauStd. Red.
AIME'240.6200.739+0.020+0.02042%
AIME'250.5330.660+0.008+0.00817%
HMMT'250.3330.635−0.022\mathbf{-0.022}16%
BrUMO'250.5880.768+0.049\mathbf{+0.049}52%
The mean effect depends on greedy - sampling alignment.

Variance reduction does not guarantee improved agreement with BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80. The greedy prior increases mean τb\tau_b on AIME'24, AIME'25, and BrUMO'25, but decreases it on HMMT'25 (table 4). At N=1N=1, when all benchmarks are pooled, this negative shift is substantially larger (table 17), indicating that an empirical prior can introduce systematic bias when greedy and sampling behave differently across datasets.

We summarize this diagnostic via greedy - sampling alignment τG-S\tau_{\text{G-S}}, defined as Kendall's τb\tau_b between the model rankings induced by greedy decoding and by stochastic sampling at N=80N=80. In our results, higher τG-S\tau_{\text{G-S}} coincides with a more positive Δτ\Delta\tau (appendix E, figure 6), suggesting that the empirical prior is most likely to help when greedy is a faithful proxy for the sampling-induced ordering. While this evidence is limited to four benchmarks, the trend is consistent with BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N acting as shrinkage toward the greedy ordering.

Model-level ranks under greedy decoding versus stochastic sampling (N=80) for each benchmark. Points on the diagonal indicate perfect alignment; color shows rank displacement ().
Figure 3. Model-level ranks under greedy decoding versus stochastic sampling (N=80N=80) for each benchmark. Points on the diagonal indicate perfect alignment; color shows rank displacement (Δ\Delta).
Implications.

BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N behaves as a shrinkage estimator toward the greedy ordering: it is helpful when greedy decoding is a faithful proxy for the sampling-induced ranking, and harmful when the two disagree. Because R0\mathbf{R}_0 is generated under a different decoding policy, incorporating it effectively biases the estimate toward greedy behavior. This can be desirable for variance reduction, but it changes the implied evaluation target. A plausible source of disagreement is that greedy decoding may under-explore on hard instances, while stochastic sampling can recover alternative successful reasoning paths. In practice, empirical priors are most attractive when NN is very small and greedy - sampling alignment has been checked on a small pilot sample; otherwise, BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N provides a safer default.

Bias - variance trade-off.

Figure 7 visualizes the trade-off induced by empirical priors: in our benchmarks, the greedy prior reduces variability (narrower distributions) but can introduce bias (shifted means), with the net effect governed by greedy - sampling alignment.

Categorical Ranking

We extend the Bayesian framework to categorical outcomes: each completion is mapped to one of C+1C+1 ordered categories based on signals such as answer format (boxed vs. unboxed), model confidence (completion bits per token), token efficiency, and external verifier judgments. Each scheme defines a categorical mapping and a utility weight vector w=(w0,…,wC)\mathbf{w}=(w_0,\dots,w_C); Bayesian estimation then proceeds with a Dirichlet - multinomial model rather than a Beta - binomial model (details and scheme definitions are given in appendix F).

We select eight non-redundant representative schemes. Using the N=1N=1 subsampling protocol on the Combined benchmark (the first L=11L=11 models of table 22, M=120M=120 questions pooled across all four datasets), we measure Kendall's τb\tau_b against three references (table 5).

Table 5. Categorical ranking at N ⁣= ⁣1N\!=\!1 on the combined benchmark (L ⁣= ⁣11L\!=\!11, the first 11 models from table 22, M ⁣= ⁣120M\!=\!120). Eight representative schemes are ordered by agreement with the gold standard (τGS\tau_{\text{GS}}, vs. BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80). Self: τb\tau_b vs. Scheme@8080; Greedy: τb\tau_b vs. BayesR0@80\mathrm{Bayes}_{\mathbf{R}_0}@80. Values are mean ±\pm std over 8080 draws.
SchemeτGS\tau_{\text{GS}}τSelf\tau_{\text{Self}}τGreedy\tau_{\text{Greedy}}
Conservative0.856±0.0760.856 \pm {\scriptstyle 0.076}0.861±0.0660.861 \pm {\scriptstyle 0.066}0.858±0.0740.858 \pm {\scriptstyle 0.074}
Efficiency-adj.0.850±0.0700.850 \pm {\scriptstyle 0.070}0.875±0.0570.875 \pm {\scriptstyle 0.057}0.859±0.0710.859 \pm {\scriptstyle 0.071}
Format-aware0.849±0.0710.849 \pm {\scriptstyle 0.071}0.881±0.0640.881 \pm {\scriptstyle 0.064}0.869±0.0690.869 \pm {\scriptstyle 0.069}
Balanced comp.0.843±0.0750.843 \pm {\scriptstyle 0.075}0.877±0.0670.877 \pm {\scriptstyle 0.067}0.862±0.0730.862 \pm {\scriptstyle 0.073}
OOD-robust0.840±0.0710.840 \pm {\scriptstyle 0.071}0.892±0.0630.892 \pm {\scriptstyle 0.063}0.870±0.0660.870 \pm {\scriptstyle 0.066}
Rare-event0.838±0.0730.838 \pm {\scriptstyle 0.073}0.888±0.0650.888 \pm {\scriptstyle 0.065}0.867±0.0690.867 \pm {\scriptstyle 0.069}
Verifier-calib.0.832±0.0760.832 \pm {\scriptstyle 0.076}0.877±0.0670.877 \pm {\scriptstyle 0.067}0.855±0.0730.855 \pm {\scriptstyle 0.073}
Verifier-only0.824±0.0710.824 \pm {\scriptstyle 0.071}0.897±0.0680.897 \pm {\scriptstyle 0.068}0.870±0.0710.870 \pm {\scriptstyle 0.071}
Self-consistency vs. gold-standard trade-off.

Signal-rich schemes achieve the highest self-consistency: Verifier-only (τSelf=0.897\tau_{\text{Self}} = 0.897) and OOD-robust (0.8920.892) rank first and second (figure 4). Yet these schemes have the lowest agreement with the gold standard (τGS=0.824\tau_{\text{GS}} = 0.824 and 0.8400.840, respectively), extending the finding from section 3.2 that high self-consistency does not imply closeness to the gold standard. The negative correlation between τGS\tau_{\text{GS}} and τSelf\tau_{\text{Self}} across schemes (figure 4) suggests that auxiliary signals introduce systematic biases away from the correctness-based ordering while stabilizing single-trial rankings.

Gold-standard agreement vs. self-consistency for 25 categorical schemes at N=1 on the Combined benchmark. Blue markers indicate the 8 representative schemes; gray markers show the remaining 17. Schemes in the upper-left are self-consistent but deviate from Bayes_ U @80 ; those in the lower-right closely track the gold standard but are less stable across single-trial draws.
Figure 4. Gold-standard agreement vs. self-consistency for 2525 categorical schemes at N=1N=1 on the Combined benchmark. Blue markers indicate the 88 representative schemes; gray markers show the remaining 1717. Schemes in the upper-left are self-consistent but deviate from BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80; those in the lower-right closely track the gold standard but are less stable across single-trial draws.
Greedy-prior alignment.

All eight schemes correlate more strongly with BayesR0@80\mathrm{Bayes}_{\mathbf{R}_0}@80 than with BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80; the gap is largest for Verifier-only (Δτ=+0.046\Delta\tau = +0.046) and OOD-robust (+0.031+0.031), consistent with the mechanism in section 3.4: verifier and OOD signals encode information partially aligned with greedy-decoding behavior. Per-dataset results (appendix F) show that scheme differentiation widens on harder benchmarks (HMMT'25, BrUMO'25), where Verifier-only drops to τGS=0.753\tau_{\text{GS}}=0.753 and 0.7340.734, while correctness-driven schemes remain stable (τGS≥0.80\tau_{\text{GS}}\ge 0.80).

Test-time scaling samples multiple solutions per prompt and aggregates them [2, 3, 4]; because stochastic reasoning varies across runs [42], we study how this variability affects rankings as budget changes. Preference evaluation and alignment learn from paired comparisons [5, 6] and underpin leaderboards such as Chatbot Arena [7, 8]. Benchmark leaderboards often rank models by task metrics such as Pass@kk [1], but item-level difficulty and discrimination affect reliability [43]; recent work adds Bayesian uncertainty and IRT-style modeling [16, 9]. We extend this literature to dense repeated-trial benchmarks and compare ranking methods by stability and convergence; appendix H gives background.

Conclusion & Future Directions

Test-time scaling turns LLM benchmarking into a repeated-sampling problem, so model rankings must be estimated from stochastic trials rather than from a single run. We formalize this setting and compare a broad collection of ranking methods within a common framework. When many trials are available, most reasonable ranking families induce nearly identical orderings, making BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N a simple and interpretable default. The main differences appear in the low-budget regime. There, uncertainty-aware estimators can improve stability, and the greedy prior BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N acts as a shrinkage estimator: it reduces variance when greedy and stochastic sampling align, but can bias rankings when they diverge.

In practice, BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N is a strong default, whereas BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N is best used after checking greedy - sampling alignment on a small pilot sample. Our experiments focus on binary correctness; extending the analysis to partial credit, rubric-based scoring, and other categorical evaluation settings is a natural next step.

Limitations

Our experiments focus on mathematical reasoning benchmarks. We do not evaluate partial credit, or open-ended outputs, where outcome categories are less clear and annotation or verification noise may be larger. More generally, when informative priors are used - especially priors derived from auxiliary signals other than greedy decoding - the prior source and specification should be reported explicitly, since the prior can introduce systematic bias if it is misaligned with the stochastic evaluation regime.

Acknowledgments

This research was supported in part by NSF awards 2117439 and 2320952.

sectionappendix appendixAppendixAppendices appendixAppendixAppendices lstlistingListingListings lstlistingListingListings

Notation and Definitions

Throughout the paper, we use the following notation.

Data and Basic Quantities

  • LL: number of models being ranked.

  • MM: number of questions in a benchmark.

  • NN: number of independent stochastic trials per model - question pair under test-time scaling.

  • R∈{0,1}L×M×N\mathbf{R}\in\{0,1\}^{L\times M\times N}: response tensor, where Rlmn=1R_{lmn}=1 if model ll solves question mm on trial nn.

  • R0\mathbf{R}_0: optional prior outcomes used by Bayesian estimators. In this paper, greedy decoding yields a shared prior matrix R0∈{0,1}M×D\mathbf{R}_0\in\{0,1\}^{M\times D} with D=1D=1, but the notation can also accommodate model-specific prior tensors.

  • p^lm:=1N∑n=1NRlmn\widehat{p}_{lm} := \frac{1}{N}\sum_{n=1}^N R_{lmn}: per-question solve rate for model ll on question mm.

  • klm:=∑n=1NRlmnk_{lm} := \sum_{n=1}^N R_{lmn}: number of successful trials for model ll on question mm.

Metric Shorthand

  • Bayes@NN: Bayesian posterior-mean estimate at NN trials under a specified prior.

  • BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N: Bayesian estimate with a uniform Dirichlet prior, denoted BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N.

  • BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N: Bayesian estimate with a greedy empirical prior, denoted BayesR0@N\mathrm{Bayes}_{\mathbf{R}_0}@N.

  • Pass@kk: probability that at least one of kk sampled completions is correct.

  • avg@NN: mean accuracy over all MM questions and NN trials. For binary outcomes, it is order-equivalent to BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N.

Ranking-Method Families

  • Pointwise methods: aggregate per-question performance to produce model scores (e.g., mean accuracy, inverse-difficulty weighting).

  • Pairwise methods: transform outcomes into win/tie counts between model pairs and fit paired-comparison models (e.g., Bradley - Terry, Elo, Glicko).

  • Listwise, setwise methods: operate on winner and loser sets for each question - trial (e.g., Plackett - Luce, Davidson - Luce).

  • Voting rules: treat questions as voters that rank models and then aggregate those preferences (e.g., Borda, Copeland, Schulze, Kemeny - Young).

  • Graph/spectral methods: construct comparison graphs and compute centrality- or flow-based scores (e.g., PageRank, Rank Centrality, HodgeRank, α\alpha-Rank).

  • IRT-inspired methods: estimate latent model abilities and item difficulties (e.g., Rasch, 2PL, 3PL, dynamic IRT).

Evaluation Criteria

  • Kendall's τb\tau_b: rank-correlation coefficient that accounts for ties; it ranges from −1-1 (perfect disagreement) to +1+1 (perfect agreement).

  • Gold-standard agreement: agreement between a low-budget ranking and the empirical gold standard, typically BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 in this paper.

  • Self-consistency: agreement between a low-budget ranking and the same method's all-trial ranking.

  • Convergence: the rate at which a method's ranking approaches its full-trial ordering as the number of trials increases.

  • Greedy - sampling alignment (τG-S\tau_{\text{G-S}}): Kendall's τb\tau_b between the ranking induced by greedy decoding and the ranking induced by stochastic sampling at high budget.

Inference Terminology

  • MLE (maximum likelihood estimation): point estimate that maximizes p(R∣θ)p(\mathbf{R}\mid\theta).

  • MAP (maximum a posteriori): point estimate that maximizes p(R∣θ)p(θ)p(\mathbf{R}\mid\theta)p(\theta).

  • EAP (expected a posteriori): posterior mean estimate E[θ∣R]\mathbb{E}[\theta\mid\mathbf{R}].

  • MML (marginal maximum likelihood): likelihood-based estimation that integrates over a latent population distribution, commonly used in IRT.

  • Credible intervals (CrI): Bayesian posterior intervals used for uncertainty quantification; we use lower credible bounds (LCBs) for conservative ranking.

Accuracy of Models

Table 6, Table 8, Table 7, Table 9 report detailed accuracy statistics for all L=20L=20 models, including greedy accuracy and stochastic-sampling statistics (minimum, mean, maximum, and standard deviation) over N=80N=80 trials. HMMT'25 is the most difficult benchmark (mean accuracies 0.0800.080 - 0.5540.554), whereas AIME'24 and BrUMO'25 are less difficult. Figure 5 visualizes these distributions across benchmarks and highlights the heterogeneity in model performance and sampling variance that motivates our ranking-stability analysis.

Table 6. Accuracy on AIME'24.
ModelGreedyTop-pp
Acc.MinMeanMaxStd
DeepSeek DS-R1-Qwen0.2000.1670.2970.4330.055
LIMO LIMO-v20.6000.4670.6190.7330.059
OpenThinker OpenThinker20.7670.6000.7220.8330.048
OpenThinker OpenThinker30.3330.4000.5170.6670.059
Qwen Qwen3-Thinking0.8670.7670.8750.9330.038
Sky-T1-32B-Flash Sky-T1-Flash0.4000.1670.3100.4000.050
gpt-oss gpt-oss-high0.7000.6330.7470.8330.053
gpt-oss gpt-oss-low0.7000.3330.6750.8670.130
gpt-oss gpt-oss-medium0.8000.5330.7550.8670.054
EXAONE EXAONE-4.00.5000.4330.5700.7330.055
NVIDIA OR-Nemotron0.4330.3670.4900.6670.064
Microsoft Phi-40.6670.5670.7050.8000.050
Microsoft Phi-4-plus0.5330.6330.7530.8670.049
OpenR1 OR1-Distill0.4000.4000.5470.7000.066
FuseO1 FuseO1-DS-QwQ-SkyT10.5000.6330.7280.8000.042
Light-R1 Light-R1-DS0.7000.6000.7340.8330.060
NVIDIA AR-Nemotron0.7000.6000.7090.8000.043
NVIDIA NVIDIA-Nemotron0.6330.5670.6760.8330.059
Qwen Qwen3-4B0.7670.6670.7720.9000.052
Bespoke-Stratos Bespoke0.1670.1000.1970.2670.043
Table 7. Accuracy on HMMT'25.
ModelGreedyTop-pp
Acc.MinMeanMaxStd
DeepSeek DS-R1-Qwen0.1330.0670.1350.2330.040
LIMO LIMO-v20.4330.2330.3470.4670.048
OpenThinker OpenThinker20.3330.2330.3820.5000.057
OpenThinker OpenThinker30.2000.2000.2970.4670.047
Qwen Qwen3-Thinking0.5000.4670.5540.6330.037
Sky-T1-32B-Flash Sky-T1-Flash0.1670.0330.1060.2000.034
gpt-oss gpt-oss-high0.2330.2670.4490.6330.069
gpt-oss gpt-oss-low0.1670.1000.2030.3330.051
gpt-oss gpt-oss-medium0.4000.3330.4550.6000.056
EXAONE EXAONE-4.00.4000.2000.3350.4330.060
NVIDIA OR-Nemotron0.2670.1670.2830.4000.049
Microsoft Phi-40.4670.2670.3780.5330.056
Microsoft Phi-4-plus0.4330.3330.4470.6330.056
OpenR1 OR1-Distill0.2330.1330.2510.3330.042
FuseO1 FuseO1-DS-QwQ-SkyT10.3000.2330.3630.4670.045
Light-R1 Light-R1-DS0.3670.2330.3560.4330.045
NVIDIA AR-Nemotron0.4000.3330.4080.5000.042
NVIDIA NVIDIA-Nemotron0.3330.2670.3620.4670.048
Qwen Qwen3-4B0.4670.3670.4640.5670.046
Bespoke-Stratos Bespoke0.0000.0000.0800.1670.035
Table 8. Accuracy on AIME'25.
ModelGreedyTop-pp
Acc.MinMeanMaxStd
DeepSeek DS-R1-Qwen0.1330.1330.2360.3330.046
LIMO LIMO-v20.6330.3330.5410.7000.068
OpenThinker OpenThinker20.5000.4670.5950.7330.060
OpenThinker OpenThinker30.3670.3330.4250.6000.057
Qwen Qwen3-Thinking0.7330.7330.8040.9000.037
Sky-T1-32B-Flash Sky-T1-Flash0.2670.1330.2200.3330.041
gpt-oss gpt-oss-high0.5670.4670.6900.8330.063
gpt-oss gpt-oss-low0.6000.2670.5980.8000.145
gpt-oss gpt-oss-medium0.5670.5000.6890.8330.065
EXAONE EXAONE-4.00.4670.3000.4410.5670.054
NVIDIA OR-Nemotron0.4330.3000.4250.5330.054
Microsoft Phi-40.6000.4000.5990.7670.072
Microsoft Phi-4-plus0.5670.5330.6830.8000.058
OpenR1 OR1-Distill0.5330.3000.4260.5670.059
FuseO1 FuseO1-DS-QwQ-SkyT10.4670.4330.5850.7330.064
Light-R1 Light-R1-DS0.6330.4670.5890.7000.056
NVIDIA AR-Nemotron0.6330.5670.6510.7330.045
NVIDIA NVIDIA-Nemotron0.4670.4330.5460.6670.050
Qwen Qwen3-4B0.7000.6000.7290.8000.044
Bespoke-Stratos Bespoke0.1000.0670.1930.3000.050
Table 9. Accuracy on BrUMO'25.
ModelGreedyTop-pp
Acc.MinMeanMaxStd
DeepSeek DS-R1-Qwen0.2670.1670.3440.5000.062
LIMO LIMO-v20.5670.5000.6510.8000.065
OpenThinker OpenThinker20.7670.6000.7380.9000.061
OpenThinker OpenThinker30.5000.4000.5120.6670.055
Qwen Qwen3-Thinking0.8670.7330.8380.9330.038
Sky-T1-32B-Flash Sky-T1-Flash0.3330.2330.3720.5000.059
gpt-oss gpt-oss-high0.4330.5330.6280.7670.053
gpt-oss gpt-oss-low0.5000.3000.3930.5000.053
gpt-oss gpt-oss-medium0.5000.5000.6100.7330.052
EXAONE EXAONE-4.00.5330.3330.4840.6330.059
NVIDIA OR-Nemotron0.4000.3330.4690.6000.054
Microsoft Phi-40.7330.5330.6920.8000.052
Microsoft Phi-4-plus0.5330.5330.7110.8000.048
OpenR1 OR1-Distill0.5670.3670.5380.6670.057
FuseO1 FuseO1-DS-QwQ-SkyT10.5670.5670.7100.9000.056
Light-R1 Light-R1-DS0.7000.6000.6900.8330.049
NVIDIA AR-Nemotron0.7670.6330.7140.8670.044
NVIDIA NVIDIA-Nemotron0.6330.5330.6490.8000.048
Qwen Qwen3-4B0.7670.6330.7440.8330.049
Bespoke-Stratos Bespoke0.1670.1670.2650.3670.053
Overview of model accuracies across all four benchmarks. Each panel shows each model's mean accuracy under stochastic sampling (over N=80 trials), together with greedy accuracy (markers). Error bars denote one standard deviation across trials and illustrate the variability introduced by test-time scaling. Models are color-coded consistently across benchmarks. The figure shows substantial heterogeneity in both absolute performance and sampling variance, with HMMT'25 notably more difficult than the other three benchmarks.
Figure 5. Overview of model accuracies across all four benchmarks. Each panel shows each model's mean accuracy under stochastic sampling (over N=80N=80 trials), together with greedy accuracy (markers). Error bars denote one standard deviation across trials and illustrate the variability introduced by test-time scaling. Models are color-coded consistently across benchmarks. The figure shows substantial heterogeneity in both absolute performance and sampling variance, with HMMT'25 notably more difficult than the other three benchmarks.

Gold Standard Agreement

To justify our use of BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 as the gold standard, we compare the full-trial rankings produced by all methods at N=80N=80. table 1 summarizes Kendall's τb\tau_b between BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 and each competing method. These results indicate that BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 is also a high-consensus ordering: by average agreement with all other methods, it ranks first on AIME'25, HMMT'25, and the Combined benchmark and second on AIME'24 and BrUMO'25 within 5×10−45\times 10^{-4} of the best (table 10). Dataset-level consensus tables appear in table 12, table 12, table 13, table 14, table 15. Many methods recover the same ordering exactly (table 16), and the remaining disagreement is concentrated in a small low-agreement tail (table 11).

Table 10. BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 as a consensus ranking. "Consensus rank" sorts methods by their average Kendall's τb\tau_b agreement with all other methods (computed at N=80N=80; ties broken by lower std).
BenchmarkMean rankBayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 avg.Best methodBest avg.Gap
AIME'2420.9414rasch_mml0.94170.0003
AIME'2510.9344avg (tie)0.93440.0000
HMMT'2510.9499avg (tie)0.94990.0000
BrUMO'2520.9542rasch_mml0.95470.0005
Combined10.9616avg (tie)0.96160.0000
Table 11. Low-agreement tail: methods whose full-trial rankings have Kendall's τb<0.85\tau_b < 0.85 relative to BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 (computed at N=80N=80).
Methodτb\tau_b
AIME'24
minimax_variant_margin_tie_ignore0.682
minimax_variant_margin_tie_half0.682
minimax_variant_winning_votes_tie_half0.682
minimax_variant_winning_votes_tie_ignore0.693
nanson_rank_ties_average0.798
nanson_rank_ties_max0.802
dynamic_irt_growth0.821
majority_judgment0.842
rasch_3pl0.842
rasch_3pl_map0.842
AIME'25
minimax_variant_winning_votes_tie_ignore0.771
majority_judgment0.779
minimax_variant_margin_tie_ignore0.819
minimax_variant_margin_tie_half0.819
minimax_variant_winning_votes_tie_half0.819
nanson_rank_ties_max0.840
nanson_rank_ties_average0.849
HMMT'25
nanson_rank_ties_max0.758
inverse_difficulty0.811
nanson_rank_ties_average0.818
minimax_variant_margin_tie_ignore0.831
minimax_variant_margin_tie_half0.831
minimax_variant_winning_votes_tie_half0.831
baldwin_rank_ties_max0.850
BrUMO'25
rasch_3pl0.789
rasch_3pl_map0.789
minimax_variant_margin_tie_ignore0.814
minimax_variant_margin_tie_half0.814
minimax_variant_winning_votes_tie_half0.814
inverse_difficulty0.821
Combined
minimax_variant_winning_votes_tie_ignore0.748
minimax_variant_margin_tie_ignore0.825
minimax_variant_margin_tie_half0.825
minimax_variant_winning_votes_tie_half0.825
nanson_rank_ties_max0.843
Table 12. Consensus ranking on AIME'24 by average Kendall's τb\tau_b agreement with all other methods at N=80N=80 (higher is better). Method variants with identical (Avg., Std.) are collapsed; we show the top 10 and bottom 5 groups.
RankMethod(s)Avg.Std.
1rasch_mml0.94170.0799
2alpharank, bayes, bayes_ci, bradley_terry_davidson, bradley_terry_davidson_map, dynamic_irt_linear, glicko_tie_draw, hodge_rank_binary_decisive, hodge_rank_binary_total, hodge_rank_binary_uniform, avg, nash_advantage_vs_equilibrium, nash_vs_equilibrium, pagerank, rank_centrality_tie_half, rasch, rasch_map, serial_rank_prob_diff, serial_rank_sign, spectral, thompson0.94140.0815
3bayes_greedy0.94070.0817
4bayesian_mcmc, bradley_terry, bradley_terry_map, elo_tie_skip, glicko_tie_correct_draw_only, glicko_tie_skip, mg_pass_at_k_2, pass_hat_k_2, plackett_luce, plackett_luce_map, trueskill0.94030.0817
5hodge_rank_log_odds_decisive, hodge_rank_log_odds_total, hodge_rank_log_odds_uniform, rank_centrality_tie_ignore, rao_kupper, rao_kupper_map0.93490.0816
6borda0.91790.0729
7baldwin_rank_ties_max0.91630.0754
8elo_tie_correct_draw_only, elo_tie_draw0.91410.0689
9copeland0.91080.0728
10schulze_tie_half0.90450.0722
Omitted ranks 11 - 22.
23nanson_rank_ties_max0.79430.0492
24nanson_rank_ties_average0.79040.0431
25dynamic_irt_growth0.78970.0763
26minimax_variant_winning_votes_tie_ignore0.68870.0691
27minimax_variant_margin_tie_half, minimax_variant_margin_tie_ignore, minimax_variant_winning_votes_tie_half0.67770.0763
RankMethod(s)Avg.Std.
1alpharank, bayes, bayes_ci, bradley_terry_davidson, bradley_terry_davidson_map, dynamic_irt_linear, hodge_rank_binary_decisive, hodge_rank_binary_total, hodge_rank_binary_uniform, avg, nash_advantage_vs_equilibrium, nash_vs_equilibrium, pagerank, rank_centrality_tie_half, rasch, rasch_map, serial_rank_prob_diff, serial_rank_sign, spectral, thompson0.93440.0595
2glicko_tie_draw0.93310.0486
3bayes_greedy0.93060.0577
4rasch_mml0.92930.0349
5hodge_rank_log_odds_decisive, hodge_rank_log_odds_total, hodge_rank_log_odds_uniform, rao_kupper, rao_kupper_map0.92850.0541
6glicko_tie_correct_draw_only0.92850.0522
7rasch_mml_credible0.92490.0216
8rasch_3pl, rasch_3pl_map0.91720.0495
9rasch_2pl, rasch_2pl_map0.91620.0507
10bradley_terry, bradley_terry_map, plackett_luce, plackett_luce_map0.91560.0486
Omitted ranks 11 - 26.
27inverse_difficulty0.84470.0455
28nanson_rank_ties_max0.83530.0256
29minimax_variant_margin_tie_half, minimax_variant_margin_tie_ignore, minimax_variant_winning_votes_tie_half0.82800.0377
30majority_judgment0.80290.0322
31minimax_variant_winning_votes_tie_ignore0.79230.0386
Table 13. Consensus ranking on HMMT'25 by average Kendall's τb\tau_b agreement with all other methods at N=80N=80 (higher is better). Method variants with identical (Avg., Std.) are collapsed; we show the top 10 and bottom 5 groups.
RankMethod(s)Avg.Std.
1alpharank, bayes, bayes_ci, bradley_terry, bradley_terry_davidson, bradley_terry_davidson_map, bradley_terry_map, dynamic_irt_linear, elo_tie_correct_draw_only, elo_tie_skip, glicko_tie_correct_draw_only, glicko_tie_draw, glicko_tie_skip, hodge_rank_binary_decisive, hodge_rank_binary_total, hodge_rank_binary_uniform, hodge_rank_log_odds_decisive, hodge_rank_log_odds_total, hodge_rank_log_odds_uniform, avg, nash_advantage_vs_equilibrium, nash_vs_equilibrium, pagerank, plackett_luce, plackett_luce_map, rank_centrality_tie_half, rao_kupper, rao_kupper_map, rasch, rasch_map, serial_rank_prob_diff, serial_rank_sign, spectral, thompson, trueskill0.94990.0631
2rasch_mml0.94940.0442
3bayes_greedy0.94680.0556
4rasch_3pl, rasch_3pl_map0.94200.0624
5rank_centrality_tie_ignore0.94150.0511
6elo_tie_draw0.93560.0561
7bayesian_mcmc0.93350.0495
8dynamic_irt_growth0.92870.0434
9pass_at_k_20.91690.0325
10rasch_2pl, rasch_2pl_map0.91610.0629
Omitted ranks 11 - 22.
23majority_judgment0.86360.0276
24minimax_variant_margin_tie_half, minimax_variant_margin_tie_ignore, minimax_variant_winning_votes_tie_half0.83910.0388
25inverse_difficulty0.83190.0458
26nanson_rank_ties_average0.81840.0156
27nanson_rank_ties_max0.77630.0346
Table 14. Consensus ranking on BrUMO'25 by average Kendall's τb\tau_b agreement with all other methods at N=80N=80 (higher is better). Method variants with identical (Avg., Std.) are collapsed; we show the top 10 and bottom 5 groups.
RankMethod(s)Avg.Std.
1rasch_mml0.95470.0581
2alpharank, bayes, bayes_ci, bayes_greedy, bradley_terry_davidson, bradley_terry_davidson_map, dynamic_irt_linear, glicko_tie_draw, hodge_rank_binary_decisive, hodge_rank_binary_total, hodge_rank_binary_uniform, hodge_rank_log_odds_decisive, hodge_rank_log_odds_total, hodge_rank_log_odds_uniform, avg, nash_advantage_vs_equilibrium, nash_vs_equilibrium, pagerank, rank_centrality_tie_half, rao_kupper, rao_kupper_map, rasch, rasch_map, serial_rank_prob_diff, serial_rank_sign, spectral, thompson0.95420.0588
3glicko_tie_correct_draw_only0.95010.0575
4mg_pass_at_k_2, pass_hat_k_20.94900.0582
5elo_tie_draw0.94230.0589
6borda, win_rate0.93990.0470
7bayesian_mcmc, bradley_terry, bradley_terry_map, elo_tie_correct_draw_only, elo_tie_skip, glicko_tie_skip, plackett_luce, plackett_luce_map, trueskill0.93820.0592
8copeland0.93760.0516
9bradley_terry_luce, bradley_terry_luce_map0.93500.0579
10dynamic_irt_growth0.93180.0505
Omitted ranks 11 - 19.
20nanson_rank_ties_average, nanson_rank_ties_max0.84370.0294
21majority_judgment0.82670.0393
22inverse_difficulty0.81560.0273
23minimax_variant_margin_tie_half, minimax_variant_margin_tie_ignore, minimax_variant_winning_votes_tie_half0.81130.0446
24rasch_3pl, rasch_3pl_map0.78930.0396
Table 15. Consensus ranking on the combined benchmark by average Kendall's τb\tau_b agreement with all other methods at N=80N=80 (higher is better). Method variants with identical (Avg., Std.) are collapsed; we show the top 10 and bottom 5 groups.
RankMethod(s)Avg.Std.
1alpharank, bayes, bayes_ci, bradley_terry_davidson, bradley_terry_davidson_map, dynamic_irt_linear, glicko_tie_draw, hodge_rank_binary_decisive, hodge_rank_binary_total, hodge_rank_binary_uniform, kemeny_young_tie_half, kemeny_young_tie_ignore, avg, nash_advantage_vs_equilibrium, nash_vs_equilibrium, pagerank, rank_centrality_tie_half, rasch, rasch_map, serial_rank_prob_diff, serial_rank_sign, spectral, thompson0.96160.0559
2copeland, ranked_pairs_strength_margin_tie_half, ranked_pairs_strength_margin_tie_ignore, ranked_pairs_strength_winning_votes_tie_half, ranked_pairs_strength_winning_votes_tie_ignore, schulze_tie_half, schulze_tie_ignore0.96020.0546
3hodge_rank_log_odds_decisive, hodge_rank_log_odds_total, hodge_rank_log_odds_uniform, rao_kupper, rao_kupper_map0.95670.0516
4glicko_tie_correct_draw_only0.95660.0522
5baldwin_rank_ties_average0.95580.0495
6borda0.95570.0516
7rank_centrality_tie_ignore0.95540.0505
8bayesian_mcmc0.95120.0493
9bayes_greedy0.95080.0473
10rasch_2pl0.95020.0486
Omitted ranks 11 - 26.
27nanson_rank_ties_average0.85620.0263
28elo_tie_draw0.85520.0239
29minimax_variant_margin_tie_half, minimax_variant_margin_tie_ignore, minimax_variant_winning_votes_tie_half0.83390.0324
30nanson_rank_ties_max0.83330.0302
31minimax_variant_winning_votes_tie_ignore0.76650.0390
Table 16. Methods that induce exactly the same ranking as BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 (τb=1\tau_b=1) when computed on the full N=80N=80 trials (excluding avg itself).
BenchmarkCountMethods
AIME'2420alpharank, bayes, bayes_ci, bradley_terry_davidson, bradley_terry_davidson_map, dynamic_irt_linear, glicko_tie_draw, hodge_rank_binary_decisive, hodge_rank_binary_total, hodge_rank_binary_uniform, nash_advantage_vs_equilibrium, nash_vs_equilibrium, pagerank, rank_centrality_tie_half, rasch, rasch_map, serial_rank_prob_diff, serial_rank_sign, spectral, thompson
AIME'2519alpharank, bayes, bayes_ci, bradley_terry_davidson, bradley_terry_davidson_map, dynamic_irt_linear, hodge_rank_binary_decisive, hodge_rank_binary_total, hodge_rank_binary_uniform, nash_advantage_vs_equilibrium, nash_vs_equilibrium, pagerank, rank_centrality_tie_half, rasch, rasch_map, serial_rank_prob_diff, serial_rank_sign, spectral, thompson
HMMT'2534alpharank, bayes, bayes_ci, bradley_terry, bradley_terry_davidson, bradley_terry_davidson_map, bradley_terry_map, dynamic_irt_linear, elo_tie_correct_draw_only, elo_tie_skip, glicko_tie_correct_draw_only, glicko_tie_draw, glicko_tie_skip, hodge_rank_binary_decisive, hodge_rank_binary_total, hodge_rank_binary_uniform, hodge_rank_log_odds_decisive, hodge_rank_log_odds_total, hodge_rank_log_odds_uniform, nash_advantage_vs_equilibrium, nash_vs_equilibrium, pagerank, plackett_luce, plackett_luce_map, rank_centrality_tie_half, rao_kupper, rao_kupper_map, rasch, rasch_map, serial_rank_prob_diff, serial_rank_sign, spectral, thompson, trueskill
BrUMO'2526alpharank, bayes, bayes_ci, bayes_greedy, bradley_terry_davidson, bradley_terry_davidson_map, dynamic_irt_linear, glicko_tie_draw, hodge_rank_binary_decisive, hodge_rank_binary_total, hodge_rank_binary_uniform, hodge_rank_log_odds_decisive, hodge_rank_log_odds_total, hodge_rank_log_odds_uniform, nash_advantage_vs_equilibrium, nash_vs_equilibrium, pagerank, rank_centrality_tie_half, rao_kupper, rao_kupper_map, rasch, rasch_map, serial_rank_prob_diff, serial_rank_sign, spectral, thompson
Combined22alpharank, bayes, bayes_ci, bradley_terry_davidson, bradley_terry_davidson_map, dynamic_irt_linear, glicko_tie_draw, hodge_rank_binary_decisive, hodge_rank_binary_total, hodge_rank_binary_uniform, kemeny_young_tie_half, kemeny_young_tie_ignore, nash_advantage_vs_equilibrium, nash_vs_equilibrium, pagerank, rank_centrality_tie_half, rasch, rasch_map, serial_rank_prob_diff, serial_rank_sign, spectral, thompson

Convergence of Ranking Methods

As the number of trials NN (or questions MM) increases, evaluation metrics such as avg@NN, Bayes@NN, and Pass@kk need not induce the same limiting ordering as ranking methods. The reason is that they target different population quantities.

We illustrate this distinction for two canonical choices used throughout the paper: the average-accuracy ranking and the Bradley - Terry (BT) model.

Large-Budget Limits: Each Method Converges, but Generally to a Different Target

To discuss M→∞M\to\infty (or N→∞N\to\infty) formally, we introduce an i.i.d. sampling model at the level of question - trial pairs. Assume (Xmn)m∈[M],n∈[N](X_{mn})_{m\in[M],n\in[N]} are i.i.d. draws from some distribution PP on {0,1}L\{0,1\}^L. Let

pℓ:=PX∼P(Xℓ=1)\begin{aligned}p_\ell &:= \mathbb{P}_{X\sim P}(X_\ell=1)\end{aligned}
wij:=PX∼P(Xi=1, Xj=0).\begin{aligned}w_{ij} &:= \mathbb{P}_{X\sim P}(X_i=1,\ X_j=0).\end{aligned}

Here, pℓp_\ell depends only on the marginal of model ℓ\ell, whereas wijw_{ij} depends on the joint distribution of (Xi,Xj)(X_i,X_j).

Average targets marginal accuracy.

By the law of large numbers,

p^ℓavg(R)  →MN→∞a.s.  pℓ.\widehat{p}^{\text{avg}}_\ell(R) \;\xrightarrow[MN\to\infty]{\text{a.s.}}\; p_\ell.

Likewise, BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N converges to the same pℓp_\ell; for binary outcomes it differs from p^ℓavg\widehat{p}^{\text{avg}}_\ell only by O((MN)−1)O((MN)^{-1}) smoothing.

Bradley - Terry targets a pairwise decisive-win functional.

The empirical win frequencies converge:

1MNWij(R)  →MN→∞a.s.  wij.\frac{1}{MN}W_{ij}(R) \;\xrightarrow[MN\to\infty]{\text{a.s.}}\; w_{ij}.

Define the BT log-likelihood

ℓ(π;W):=∑i≠jWij(log⁡πi−log⁡(πi+πj)).\ell(\pi;W) := \sum_{i\neq j} W_{ij}\Bigl(\log \pi_i - \log(\pi_i+\pi_j)\Bigr). (8)

Then the BT-ML estimator is an MM-estimator: maximizing (8) with WijW_{ij} is equivalent to maximizing the scaled objective (MN)−1ℓ(π;W)(MN)^{-1}\ell(\pi;W). Under mild regularity and connectivity conditions (ensuring strict concavity in log⁡π\log\pi and uniqueness up to scale), π^\widehat{\pi}converges to the unique (up to scale) maximizer of the population objective

π⋆∈arg⁡max⁡π>0∑i≠jwij(log⁡πi−log⁡(πi+πj)).\pi^\star \in \arg\max_{\pi>0} \sum_{i\neq j} w_{ij}\Bigl(\log \pi_i - \log(\pi_i+\pi_j)\Bigr). (9)

The limiting objects (pℓ)ℓ=1L(p_\ell)_{\ell=1}^L and π⋆\pi^\star are generally not linked by any monotone transform: pℓp_\ell depends only on marginal correctness, while π⋆\pi^\star depends on the full matrix (wij)i≠j(w_{ij})_{i\neq j}. Therefore, without additional assumptions on PP (e.g., that PP is generated by a BT choice model at the level of decisive comparisons), there is no reason to expect the induced orderings to coincide as MN→∞MN\to\infty. The following counterexample demonstrates this non-equivalence.

A Counterexample: Average accuracy and Bradley - Terry disagree even at infinite budget

We construct a distribution PP (equivalently, a finite pattern that can be repeated) for which the average ranking and the BT-ML ranking disagree. The construction uses L=3L=3 models. For notational convenience, we label them 0,1,20,1,2.

Outcome patterns.

Consider the following three outcome vectors in {0,1}3\{0,1\}^3:

Type A: (0,1,1),\begin{aligned}\text{Type A: } &(0,1,1),\end{aligned}
Type B: (1,0,0),\begin{aligned}\text{Type B: } &(1,0,0),\end{aligned}
Type C: (1,1,0).\begin{aligned}\text{Type C: } &(1,1,0).\end{aligned}

Let PP place mass

P(A)=28,P(B)=38,P(C)=38.\mathbb{P}(\text{A})=\frac{2}{8},\qquad \mathbb{P}(\text{B})=\frac{3}{8},\qquad \mathbb{P}(\text{C})=\frac{3}{8}.

Equivalently, one may take a deterministic dataset with M=8M=8 questions and N=1N=1 trial, containing exactly 22 questions of Type A, 33 of Type B, and 33 of Type C; repeating this block preserves both rankings, as established by the derivation.

The marginal success probabilities are

p0=68=34,p1=58,p2=28=14,p_0 = \frac{6}{8}=\frac34,\qquad p_1 = \frac{5}{8},\qquad p_2 = \frac{2}{8}=\frac14,

so the average method ranks

0  >  1  >  2.0 \;>\; 1 \;>\; 2.

For these three types, the decisive-win probabilities wij=P(Xi=1,Xj=0)w_{ij}=\mathbb{P}(X_i=1,X_j=0) are:

w01=38,w10=28,w02=68,w20=28,w12=38,w21=0.\begin{aligned} w_{01}&=\frac{3}{8},\quad w_{10}=\frac{2}{8},\\ w_{02}&=\frac{6}{8},\quad w_{20}=\frac{2}{8},\\ w_{12}&=\frac{3}{8},\quad w_{21}=0. \end{aligned}

For the finite M=8,N=1M=8,N=1 realization, the corresponding win counts are Wij=8wijW_{ij}=8w_{ij}, i.e.,

W=(036203200).W = \begin{pmatrix} 0 & 3 & 6 \\ 2 & 0 & 3 \\ 2 & 0 & 0 \end{pmatrix}. (10)

It remains to show that BT-ML ranks 1>0>21>0>2 for (10), thereby disagreeing with the average ranking.

A convenient characterization of the BT-ML optimum is the standard first-order condition equating observed wins to model-implied expected wins: for each ii,

∑j≠iWij  =  ∑j≠i(Wij+Wji)⋅πiπi+πj.\sum_{j\neq i} W_{ij} \;=\; \sum_{j\neq i} (W_{ij}+W_{ji})\cdot \frac{\pi_i}{\pi_i+\pi_j}. (11)

(These equations follow by differentiating (8) with respect to log⁡πi\log\pi_i.)

Because BT strengths are identifiable only up to a global scale factor, fix π2=1\pi_2=1 and write π0=a\pi_0=a, π1=b\pi_1=b. Plugging (10) into (11) yields two independent equations:

9=5⋅aa+b  +  8⋅aa+1,\begin{aligned}9 &= 5\cdot \frac{a}{a+b} \;+\; 8\cdot \frac{a}{a+1}, \end{aligned} (12)
5=5⋅ba+b  +  3⋅bb+1.\begin{aligned}5 &= 5\cdot \frac{b}{a+b} \;+\; 3\cdot \frac{b}{b+1}. \end{aligned} (13)
Step 1: solve aa in terms of bb.

From (13),

5−3⋅bb+1=5⋅ba+b.5 - 3\cdot \frac{b}{b+1} = 5\cdot \frac{b}{a+b}.

The left-hand side simplifies:

5−3⋅bb+1=5(b+1)−3bb+1=2b+5b+1.5 - 3\cdot\frac{b}{b+1} = \frac{5(b+1)-3b}{b+1} = \frac{2b+5}{b+1}.

Thus,

5ba+b=2b+5b+1\begin{aligned}\frac{5b}{a+b} &= \frac{2b+5}{b+1}\end{aligned} (14)
⟹a+b=5b(b+1)2b+5\begin{aligned}&\quad\Longrightarrow\quad a+b = \frac{5b(b+1)}{2b+5}\end{aligned} (15)
⟹a=3b22b+5.\begin{aligned}&\quad\Longrightarrow\quad a = \frac{3b^2}{2b+5}. \end{aligned} (16)
Step 2: determine bb from a one-dimensional equation.

Substitute (16) into (12). The substitution gives

aa+b=3b22b+55b(b+1)2b+5=3b5(b+1),\frac{a}{a+b} = \frac{\frac{3b^2}{2b+5}}{\frac{5b(b+1)}{2b+5}} = \frac{3b}{5(b+1)},

so the first term in (12) becomes 5⋅aa+b=3bb+15\cdot \frac{a}{a+b} = \frac{3b}{b+1}. Also,

aa+1=3b22b+53b22b+5+1=3b23b2+2b+5.\frac{a}{a+1} = \frac{\frac{3b^2}{2b+5}}{\frac{3b^2}{2b+5}+1} = \frac{3b^2}{3b^2+2b+5}.

Therefore (12) is equivalent to

9=3bb+1+8⋅3b23b2+2b+5.9 = \frac{3b}{b+1} + 8\cdot \frac{3b^2}{3b^2+2b+5}.

Simplifying gives the cubic equation

2b3−5b2−16b−15=0.2b^3 - 5b^2 - 16b - 15 = 0. (17)

Let f(b)=2b3−5b2−16b−15f(b)=2b^3 - 5b^2 - 16b - 15. We have

f(4)=128−80−64−15=−31<0,\begin{aligned}f(4) &= 128-80-64-15 = -31 < 0,\end{aligned}
f(5)=250−125−80−15=30>0,\begin{aligned}f(5) &= 250-125-80-15 = 30 > 0,\end{aligned}

so there exists a root b⋆∈(4,5)b^\star\in(4,5). Moreover,

f′(b)=6b2−10b−16=2(3b2−5b−8),f'(b)=6b^2-10b-16 = 2(3b^2-5b-8),

whose positive root is b=5+1216=166=83b=\frac{5+\sqrt{121}}{6}=\frac{16}{6}=\frac{8}{3}. Hence ff is strictly increasing for all b>83b>\frac{8}{3}, implying the root b⋆∈(4,5)b^\star\in(4,5) is unique. We therefore conclude that the BT-ML solution (under π2=1\pi_2=1) satisfies b=b⋆∈(4,5)b=b^\star\in(4,5) and a=3(b⋆)22b⋆+5a=\frac{3(b^\star)^2}{2b^\star+5}.

Step 3: show b>a>1b>a>1, hence BT ranks 1>0>21>0>2.

Using (16),

ba=b3b22b+5=2b+53b.\frac{b}{a} = \frac{b}{\frac{3b^2}{2b+5}} = \frac{2b+5}{3b}.

Thus, b>ab>a holds exactly when (2b+5)/(3b)>1(2b+5)/(3b)>1, i.e., when b<5b<5. Since b⋆∈(4,5)b^\star\in(4,5), we have b⋆>ab^\star>a.

It remains to show a>1a>1. Suppose for contradiction that a≤1a\le 1. We have already established b>ab>a, so b≥ab\ge a. Then aa+b≤a2a=12\frac{a}{a+b}\le \frac{a}{2a}=\frac12 and aa+1≤12\frac{a}{a+1}\le \frac12. Plugging into (12) gives

9=5⋅aa+b+8⋅aa+1≤5⋅12+8⋅12=132<9,9 = 5\cdot \frac{a}{a+b} + 8\cdot \frac{a}{a+1} \le 5\cdot\frac12 + 8\cdot\frac12 = \frac{13}{2} < 9,

a contradiction. Hence a>1=π2a>1=\pi_2. Putting these inequalities together yields

π1=b⋆  >  π0=a  >  π2=1,\pi_1=b^\star \;>\; \pi_0=a \;>\; \pi_2=1,

so BT ranks

1  >  0  >  2.1 \;>\; 0 \;>\; 2.

This contradicts the average ranking 0>1>20>1>2, establishing that the two methods can induce different orderings even in the absence of sampling noise.

From a Finite Counterexample to "No Convergence" as M→∞M\to\infty or N→∞N\to\infty.

The counterexample rules out a general theorem forcing average and BT rankings to coincide in the large-budget limit. To connect it directly to M→∞M\to\infty or N→∞N\to\infty, it suffices that both methods are invariant under replication.

Replication invariance (deterministic construction).

Let RR be any fixed tensor. For an integer k≥1k\ge 1, define: (i) question replication R(k,M)R^{(k,M)} by repeating the MM questions kk times (so M′=kMM' = kM and N′=NN' = N), and (ii) trial replication R(k,N)R^{(k,N)} by repeating the NN trials kk times (so M′=MM'=M and N′=kNN' = kN). Then:

  1. Average scores are unchanged:

    p^ℓavg(R(k,M))=p^ℓavg(R(k,N))=p^ℓavg(R).\widehat{p}^{\text{avg}}_\ell(R^{(k,M)})=\widehat{p}^{\text{avg}}_\ell(R^{(k,N)})=\widehat{p}^{\text{avg}}_\ell(R).
  2. The decisive-win matrix scales linearly:

    W(R(k,M))=k W(R)\begin{aligned}W(R^{(k,M)}) &= k\,W(R)\end{aligned}
    W(R(k,N))=k W(R).\begin{aligned}W(R^{(k,N)}) &= k\,W(R).\end{aligned}
  3. The BT-ML maximizer is unchanged, because the log-likelihood scales as

    ℓ(π;kW)=k ℓ(π;W),\ell(\pi; kW) = k\,\ell(\pi;W),

    and therefore has the same maximizer.

Therefore, if two methods disagree on RR, they disagree on R(k,M)R^{(k,M)} for arbitrarily large MM and on R(k,N)R^{(k,N)} for arbitrarily large NN. Applied to the M=8,N=1M=8,N=1 tensor corresponding to (10), this yields an explicit sequence with M→∞M\to\infty (or N→∞N\to\infty) for which the average and BT rankings remain different at every budget.

Stochastic formulation (i.i.d. construction).

Alternatively, under the i.i.d. model of appendix C.2, the same discrepancy appears at the population level. For the distribution PP in appendix C.2.0, the limiting average ranking is determined by (pℓ)(p_\ell) and yields 0>1>20>1>2, while the limiting BT ranking is determined by the maximizer of (9) and yields 1>0>21>0>2. Thus, even with independent sampling and MN→∞MN\to\infty, the two rankings can converge to different limits.

Implications and Support for the Gold-Standard Definition

This analysis has a direct implication for benchmarking ranking methods: there is no method-independent guarantee that all reasonable procedures converge to the same ordering as the evaluation budget grows. Different ranking procedures correspond to different statistical targets.

Why this happens.

Average-based ranking targets the marginal success probabilities pℓ=P(Xℓ=1)p_\ell=\mathbb{P}(X_\ell=1). BT instead targets the latent strengths that best explain the decisive pairwise win rates wij=P(Xi=1,Xj=0)w_{ij}=\mathbb{P}(X_i=1,X_j=0) through a logistic choice model. These are different summaries of the same joint outcome distribution PP. The counterexample in appendix C.2.0 isolates the mechanism: a model can have higher marginal accuracy while assigning less decisive-win mass against another model, which shifts the BT optimum.

Why a gold standard is needed.

Because ranking methods need not share a common asymptotic ordering, claims about "distance to the truth" require a specified target ordering. Otherwise even statements such as "method AA converges faster than method BB" are ambiguous.

Our choice: BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N.

We define the gold-standard ordering as BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N (with N=80N=80 in our experiments). This definition is supported by three considerations:

  1. Interpretability and decision relevance. BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N estimates the probability that a model solves a randomly drawn benchmark item under the sampling policy. This is an accuracy-like quantity with a direct operational meaning.

  2. Minimal modeling assumptions. BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N (and avg@NN) depend only on marginal correctness and do not impose a parametric pairwise-choice model. Methods such as BT are useful when the pairwise-choice model is appropriate, but their induced ordering is not, in general, a refinement of accuracy.

  3. Consistency under increasing budget. Under i.i.d. sampling of (m,n)(m,n) pairs, BayesU@N\mathrm{Bayes}_{\mathcal{U}}@N converges to pℓp_\ell as MN→∞MN\to\infty, making it a natural "infinite-budget" reference for accuracy-based evaluation.

Relationship to self-consistency.

This non-convergence result does not argue against BT or other rankers. It instead clarifies that two evaluations are complementary: agreement with an explicit accuracy-based target, and self-consistency, i.e., how quickly a method stabilizes toward its own full-budget ordering. The former asks whether a method matches the chosen reference; the latter asks how stable the method itself becomes as trials accumulate. The counterexample shows why these questions are not interchangeable.

Minimality of the eight-question construction

The counterexample in Section C.2.0 uses M=8M=8 questions. The same setting (L=3L=3, N=1N=1, and BT-ML fit from decisive wins) also yields a minimality fact: there is no strict disagreement example with fewer than eight questions.

Proposition (minimal MM for strict disagreement; verified by exhaustive enumeration).

Assume L=3L=3 and N=1N=1. Assume moreover that the average ranking is strict (all three average scores are distinct), and that BT-ML is well-defined and finite (equivalently, the directed win graph with an edge i→ji\to j whenever Wij>0W_{ij}>0 is strongly connected, which ensures a unique BT-ML maximizer up to global scale). If BT-ML disagrees with the average ranking, then M≥8M\ge 8.

Verification.

With N=1N=1, each question produces an outcome pattern in {0,1}3\{0,1\}^3. Hence, up to permutation of questions, any dataset with MM questions is determined by the count vector c=(cx)x∈{0,1}3∈N8c=(c_x)_{x\in\{0,1\}^3}\in\mathbb{N}^8 with ∑xcx=M\sum_x c_x=M. For fixed MM, there are (M+77)\binom{M+7}{7} such vectors; thus the total number of datasets with M≤7M\le 7 is

∑M=17(M+77)  =  6434.\sum_{M=1}^7 \binom{M+7}{7} \;=\; 6434.

For each such dataset, we compute the induced average ordering and the BT-ML ordering (obtained by maximizing (8), equivalently solving (11)). Restricting to datasets with (i) strict average ordering and (ii) strong connectivity (so the BT-ML maximizer is unique up to scale), an exhaustive enumeration yields 15061506 instances for M≤7M\le 7; in all of them the BT-ML ordering agrees with the average ordering. Therefore, no strict-disagreement example exists for M≤7M\le 7.

Section C.2.0 exhibits a strict-disagreement dataset at M=8M=8, so Mmin=8M_{\mathrm{min}}=8.

Ranking-Method Stability at N=1N=1

We provide additional details for the N=1N=1 stability analyses in section 3.2. Method rankings on the Combined benchmark are reported for (i) gold-standard agreement (method@1 vs. BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80) and (ii) self-consistency (method@1 vs. method@80), collapsing method variants with identical mean and standard deviation across the 80 single-trial draws.

As a pair-level diagnostic, we compute gap-conditional stability: pooled over benchmarks, BayesU@1\mathrm{Bayes}_{\mathcal{U}}@1 reversals concentrate among near-tied pairs under the BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 gap gij=∣μi−μj∣g_{ij}=|\mu_i-\mu_j| (gij≤0.02g_{ij}\le0.02: reversal 0.3500.350, tie 0.1510.151, pairwise correctness 0.5660.566), while well-separated pairs are ordered almost perfectly (gij≥0.10g_{ij}\ge0.10: correctness 0.9850.985; gij≥0.20g_{ij}\ge0.20: correctness 0.9990.999).

Table 17. Gold-standard agreement at N=1N=1 on the combined benchmark, measured as Kendall's τb\tau_b between each method's single-trial ranking and the gold standard (BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80). Statistics are computed over 8080 single-trial draws. Methods with identical mean/std. values are collapsed; we show the top 10 and bottom 5 groups.
RankMethod(s)MeanStd.
1baldwin_rank_ties_average, bayes, bayes_ci, borda, copeland, majority_judgment, avg, minimax_variant_margin_tie_half, minimax_variant_margin_tie_ignore, minimax_variant_winning_votes_tie_half, nash_advantage_vs_equilibrium, nash_vs_equilibrium, pagerank, rank_centrality_tie_half, ranked_pairs_strength_margin_tie_half, ranked_pairs_strength_margin_tie_ignore, ranked_pairs_strength_winning_votes_tie_half, ranked_pairs_strength_winning_votes_tie_ignore, schulze_tie_half, schulze_tie_ignore, spectral0.86470.0486
2alpharank0.86460.0486
3rasch_mml_credible0.86420.0351
4hodge_rank_binary_uniform0.86230.0491
5hodge_rank_binary_decisive0.86230.0484
6hodge_rank_binary_total0.86160.0493
7serial_rank_sign0.86150.0503
8hodge_rank_log_odds_total, hodge_rank_log_odds_uniform0.86030.0482
9rao_kupper_map0.86030.0483
10rao_kupper0.86010.0484
Omitted ranks 11 - 38.
39nanson_rank_ties_average0.80670.0363
40bradley_terry_luce_map0.80640.0556
41bradley_terry_luce0.80580.0554
42bayes_greedy0.78560.0309
43nanson_rank_ties_max0.78250.0394
Table 18. Self-consistency at N=1N=1 on the combined benchmark, measured as Kendall's τb\tau_b between each method's single-trial ranking and its own full-trial ranking (method@80). Statistics are computed over 80 single-trial draws. Methods with identical (Mean, Std.) are collapsed; we show the top 10 and bottom 5 groups.
RankMethod(s)MeanStd.
1nanson_rank_ties_average0.89250.0497
2rasch_mml_credible0.88310.0370
3nanson_rank_ties_max0.86690.0589
4baldwin_rank_ties_average0.86640.0492
5copeland, ranked_pairs_strength_margin_tie_half, ranked_pairs_strength_margin_tie_ignore, ranked_pairs_strength_winning_votes_tie_half, ranked_pairs_strength_winning_votes_tie_ignore, schulze_tie_half, schulze_tie_ignore0.86540.0489
6rasch_mml0.86480.0417
7bayes, bayes_ci, avg, nash_advantage_vs_equilibrium, nash_vs_equilibrium, pagerank, rank_centrality_tie_half, spectral0.86470.0486
8alpharank0.86460.0486
9borda0.86460.0499
10hodge_rank_binary_uniform0.86230.0491
Omitted ranks 11 - 44.
45elo_tie_correct_draw_only0.80740.0507
46bayes_greedy0.80640.0309
47elo_tie_draw0.80630.0507
48minimax_variant_margin_tie_half, minimax_variant_margin_tie_ignore, minimax_variant_winning_votes_tie_half0.79630.0454
49minimax_variant_winning_votes_tie_ignore0.76550.0455

Additional Prior Diagnostics

Supplementary diagnostics for the empirical-prior analysis in section 3.4 are shown in figure 6, figure 7.

Across our four benchmarks, the prior advantage is not monotonically related to difficulty (a), but it is associated with greedy - sampling alignment (b). The sampling - greedy accuracy gap (c) shows no clear relationship.
Figure 6. Across our four benchmarks, the prior advantage is not monotonically related to difficulty (a), but it is associated with greedy - sampling alignment (b). The sampling - greedy accuracy gap (c) shows no clear relationship.
Bootstrap distributions of Kendall's _b at N=1 (50 samples). Violin plots show the full distribution; the greedy prior (red) yields narrower distributions but can shift the mean negatively (HMMT'25 ) or positively (BrUMO'25 ).
Figure 7. Bootstrap distributions of Kendall's τb\tau_b at N=1N=1 (5050 samples). Violin plots show the full distribution; the greedy prior (red) yields narrower distributions but can shift the mean negatively (HMMT'25) or positively (BrUMO'25).

Categorical Ranking

We report the experimental setup and per-dataset results for the categorical-ranking experiments summarized in section 3.5.

Setup

The binary Bayesian estimator (section 2.3) models each trial outcome as Rlmn∈{0,1}R_{lmn}\in\{0,1\} and places a Beta prior on the per-question solve rate. The categorical extension maps each completion to one of C+1C+1 categories, yielding outcomes Rlmn∈{0,…,C}R_{lmn}\in\{0,\dots,C\} defined by auxiliary signals extracted during generation. A categorical scheme ss specifies:

  1. a categorical mapping ϕs ⁣:completion features→{0,…,Cs}\phi_s\colon \text{completion features}\to\{0,\dots,C_s\}, which assigns each completion to a category based on predicates over the base signals (table 19), and

  2. a utility weight vector ws=(w0,…,wCs)∈RCs+1\mathbf{w}_s=(w_0,\dots,w_{C_s})\in\mathbb{R}^{C_s+1}, encoding the relative value of each category.

Bayesian estimation replaces the Beta - binomial model with a Dirichlet - multinomial model: for each model - question pair, a symmetric Dirichlet prior is placed on the C+1C+1 category probabilities θ=(θ0,…,θC)\boldsymbol{\theta}=(\theta_0,\dots,\theta_C), and the posterior mean of the weighted utility ∑k=0Cwkθ^k\sum_{k=0}^{C}w_k\hat{\theta}_k is computed. Model-level scores are then aggregated across questions, as in the binary case.

Base signals.

For each of the L=11L=11 models in the categorical cohort, we extract 99 features per completion (table 19). These features span five domains: answer format (has_box), correctness (is_correct), generation cost (token_ratio, repeated_pattern), decoding confidence (prompt_bpt, completion_bpt), and external verification via CompassVerifier (compass_A/B/C). The feature tensors have shape (N,M,9)(N, M, 9) per model, with N=80N=80 trials and M=30M=30 questions per benchmark.

CompassVerifier CompassVerifier-3B provides the external verification signals. Its scores on completions generated by the other models define the verifier-based categorical schemes. Verifier inference uses Transformers [44] and Accelerate [45], with FlashAttention kernels [46] and the DFloat11 format [47] for throughput.

Table 19. Nine base signals extracted per completion for the categorical ranking experiments. Each model - question - trial entry produces a vector in R9\mathbb{R}^9.
#SignalDescription
1has_boxBoxed final answer present (0/1)
2is_correctExact-match correctness (0/1)
3token_ratioCompletion tokens / 32768
4repeated_patternNon-stop finish reason (0/1)
5prompt_bptPrompt bits-per-token
6completion_bptCompletion bits-per-token
7compass_AVerifier P(correct)P(\text{correct})
8compass_BVerifier P(wrong)P(\text{wrong})
9compass_CVerifier P(irrelevant)P(\text{irrelevant})
Derived predicates and thresholds.

Several predicates are shared across schemes. All thresholds are computed per-model from the available samples:

  • Invalid: repeated_pattern=1\texttt{repeated\_pattern}=1 or compass_C≥0.5\texttt{compass\_C}\ge 0.5.

  • Confidence: High confidence :=:= completion_bpt≤P40(completion_bpt)\texttt{completion\_bpt}\le P_{40}(\texttt{completion\_bpt}); wrong - high-confidence :=:= wrong and completion_bpt≤P60(completion_bpt∣wrong)\texttt{completion\_bpt}\le P_{60}(\texttt{completion\_bpt}\mid\text{wrong}).

  • Prompt OOD: prompt_bpt≥P90(prompt_bpt)\texttt{prompt\_bpt}\ge P_{90}(\texttt{prompt\_bpt}).

  • Efficiency bands: Economical/moderate/verbose based on P33P_{33} and P66P_{66} of token_ratio\texttt{token\_ratio}.

  • Verifier: CompassVerifier dominant label is arg⁡max⁡(A,B,C)\arg\max(A,B,C); verifier-high :=:= A≥0.6A\ge 0.6.

Scheme definitions.

The experiments include 2525 categorical schemes spanning correctness-only baselines (A, H, S, Y), confidence-aware (C, I, J, V), format-aware (B, P, T), efficiency-aware (F, G, M), verifier-based (D, K, O, U, Z), OOD-aware (E, N, W), abstention-aware (L, Q), and composite (R) variants. Several schemes are metric-level near-duplicates (e.g., A ≡\equiv S, H ≡\equiv Y, L ≡\equiv Q), so the reported comparison uses 88 non-redundant representative schemes covering distinct design axes (table 20).

Table 20. Eight representative categorical schemes used in section 3.5. Each scheme maps a completion to one of C+1C+1 categories using the base signals in table 19 and scores the result with a utility weight vector w\mathbf{w}. Category 00 is always Invalid (w0=0w_0=0) unless otherwise noted.
SchemeIntentCategories (kk)Weights w\mathbf{w}
ConservativePenalize confidently-wrong1: Wrong ∧\wedge HighConf
2: Wrong ∧\wedge LowConf
3: Correct
(0, −0.10, 0.05, 1.00)(0,\, {-}0.10,\, 0.05,\, 1.00)
Efficiency-adj.Discount verbose correct1 - 3: Wrong ×\times {Econ., Mod., Verb.}
4 - 6: Correct ×\times {Econ., Mod., Verb.}
(0, 0.10, 0.07, 0.03, 1, 0.92, 0.85)(0,\, 0.10,\, 0.07,\, 0.03,\, 1,\, 0.92,\, 0.85)
Format-awareReward boxed correct1: Wrong ∧\wedge Unboxed; 2: Wrong ∧\wedge Boxed
3: Correct; 4: Correct ∧\wedge Boxed
(0, 0.10, 0.05, 0.90, 1)(0,\, 0.10,\, 0.05,\, 0.90,\, 1)
Balanced comp.Format ×\times confidence1: Wrong ∧\wedge Unboxed; 2 - 3: Wrong ∧\wedge Boxed ×\times Conf
4 - 7: Correct ×\times {Un/Boxed} ×\times Conf
(0, 0.10, 0.06, −0.02, 0.90, 0.95, 0.97, 1)(0,\, 0.10,\, 0.06,\, {-}0.02,\, 0.90,\, 0.95,\, 0.97,\, 1)
OOD-robustReward in-distribution1: OOD ∧\wedge Wrong; 2: InDist ∧\wedge Wrong
3: OOD ∧\wedge Correct; 4: InDist ∧\wedge Correct
(0, 0.05, 0.10, 0.95, 1)(0,\, 0.05,\, 0.10,\, 0.95,\, 1)
Rare-eventOOD + abstention1: OOD ∧\wedge Wrong; 2: OOD ∧\wedge Correct
3: InDist ∧\wedge Wrong; 4: InDist ∧\wedge Correct; 5: Abstain
(0, 0.05, 1, 0.08, 0.95, 0.20)(0,\, 0.05,\, 1,\, 0.08,\, 0.95,\, 0.20)
Verifier-calib.Penalize false-positive1: Wrong ∧\wedge A≥0.6A{\ge}0.6; 2: Wrong ∧\wedge A<0.6A{<}0.6
3 - 5: Correct ×\times {AlowA_\text{low}, AmidA_\text{mid}, AhighA_\text{high}}
(0, −0.05, 0.05, 0.88, 0.94, 1)(0,\, {-}0.05,\, 0.05,\, 0.88,\, 0.94,\, 1)
Verifier-onlyNo ground truth0: Repeated; 1: Dominant =C=C
2: Dominant =B=B; 3: Dominant =A=A
(0, 0, 0.1, 1)(0,\, 0,\, 0.1,\, 1)
Evaluation protocol.

For each scheme, we apply the same N=1N=1 subsampling protocol as in section 3.2: one of the N=80N=80 trials is subsampled per question, the scheme's categorical ranking is computed, and Kendall's τb\tau_b is measured against three references:

  1. Gold-standard (τGS\tau_{\text{GS}}): agreement with the binary BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 ranking, which treats outcomes as correct/wrong with a uniform Dirichlet prior.

  2. Self-consistency (τSelf\tau_{\text{Self}}): agreement with the scheme's own all-8080-trial ranking (Scheme@8080).

  3. Greedy-prior (τGreedy\tau_{\text{Greedy}}): agreement with BayesR0@80\mathrm{Bayes}_{\mathbf{R}_0}@80, the binary Bayes ranking incorporating a greedy-decoding empirical prior.

Statistics (mean and standard deviation) are computed over the 8080 single-trial draws. Combined results aggregate the four benchmarks (M=120M=120 questions) and are reported in table 5; per-dataset results are reported in table 21.

Per-Dataset Results

table 21 reports gold-standard agreement and self-consistency for each benchmark separately. The results show three main patterns.

Table 21. Per-dataset categorical ranking at N ⁣= ⁣1N\!=\!1. Gold-standard agreement (τGS\tau_{\text{GS}}: vs. BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80) and self-consistency (τSelf\tau_{\text{Self}}: vs. Scheme@8080) for the 88 representative schemes. Values are mean Kendall's τb\tau_b over 8080 single-trial draws.
Gold-standard agreement (τGS\tau_{\text{GS}})Self-consistency (τSelf\tau_{\text{Self}})
SchemeAIME'24AIME'25HMMT'25BrUMO'25AIME'24AIME'25HMMT'25BrUMO'25
Conservative0.8140.8130.8010.8150.8140.8130.8010.820
Efficiency-adj.0.8140.8210.8120.8170.8140.8140.8140.828
Format-aware0.8200.8080.8120.8190.8200.8110.8130.830
Balanced comp.0.8160.8100.8040.8060.8160.8130.8050.816
OOD-robust0.8190.8060.7880.8100.8190.8030.8020.824
Rare-event0.8160.8040.7930.8160.8160.8010.8070.828
Verifier-calib.0.8170.8020.7860.7960.8100.8010.8000.807
Verifier-only0.8130.8050.7530.7340.8060.8090.7950.810
Narrow spread on individual benchmarks.

On each benchmark individually, all eight schemes achieve τGS\tau_{\text{GS}} between 0.730.73 and 0.830.83, with inter-scheme variation much smaller than on the combined benchmark. On AIME'24, the range across the 88 schemes is only 0.0070.007 (0.8130.813 - 0.8200.820). This narrow spread reflects the limited information available from a single trial with M=30M=30 questions and L=11L=11 models; the combined benchmark (M=120M=120) offers finer discrimination among category structures.

Verifier-only degrades on hard benchmarks.

The Verifier-only scheme exhibits the largest performance drop on the harder benchmarks: τGS\tau_{\text{GS}} falls from 0.8130.813 (AIME'24) to 0.7530.753 (HMMT'25) and 0.7340.734 (BrUMO'25), a decline of 0.060.06 - 0.080.08. In contrast, correctness-driven schemes (Conservative, Efficiency-adjusted, Format-aware) remain above 0.800.80 on all benchmarks. This pattern suggests that CompassVerifier judgments are less reliable proxies for correctness on more challenging problems.

Self-consistency converges to gold-standard on individual benchmarks.

On AIME'24, the self-consistency column is nearly identical to the gold-standard column for most schemes, indicating that the all-8080 scheme ranking coincides with the binary BayesU@80\mathrm{Bayes}_{\mathcal{U}}@80 ranking when the number of questions is small. On the combined benchmark (table 5), self-consistency consistently exceeds gold-standard agreement, reflecting convergence of each scheme to its own distinct ordering when given enough questions.

Arena Ranking

Our experiments consider a dense benchmark tensor R∈{0,1}L×M×N\mathbf{R}\in\{0,1\}^{L\times M\times N}: every model is evaluated on every question, and repeated stochastic trials provide multiple binary outcomes for each model - question pair. Arena evaluation, exemplified by preference leaderboards such as Chatbot Arena, uses a different observation model. Its primitive datum is a comparison among a small set of model responses to a prompt, so the data form a sparse, possibly time-varying comparison graph rather than a complete model - question - trial tensor. We analyze how the ranking families studied in this work transfer to this sparse-comparison regime, and where additional assumptions are required.

Observation Model

Let

Darena={(at,bt,xt,yt,ωt)}t=1T\mathcal{D}_{\mathrm{arena}} =\{(a_t,b_t,x_t,y_t,\omega_t)\}_{t=1}^{T}

denote a pairwise Arena log. At comparison tt, models at,bt∈{1,…,L}a_t,b_t\in\{1,\dots,L\} respond to prompt xtx_t; a human or model judge returns yt∈{at≻bt, bt≻at, at∼bt}y_t\in\{a_t\succ b_t,\ b_t\succ a_t,\ a_t\sim b_t\}; and ωt≥0\omega_t\ge0 is an optional weight for reliability, deduplication, or target-distribution reweighting. Writing i≻tji\succ_t j for the event that model ii is preferred to model jj in comparison tt, let Tij={t:{at,bt}={i,j}}\mathcal{T}_{ij}=\{t:\{a_t,b_t\}=\{i,j\}\}. We define

Wij=∑t∈Tijωt1{i≻tj},W_{ij}=\sum_{t\in\mathcal{T}_{ij}}\omega_t\mathbf{1}\{i\succ_t j\},
Tij=∑t∈Tijωt1{at∼bt},T_{ij}=\sum_{t\in\mathcal{T}_{ij}}\omega_t\mathbf{1}\{a_t\sim b_t\},

with Wii=Tii=0W_{ii}=T_{ii}=0 and Cij=Wij+Wji+TijC_{ij}=W_{ij}+W_{ji}+T_{ij}. These counts induce the comparison graph

Garena=(V,E),V={1,…,L},G_{\mathrm{arena}}=(V,E),\qquad V=\{1,\dots,L\},
E={{i,j}:Cij>0}.E=\{\{i,j\}:C_{ij}>0\}.

In the dense benchmark setting, each question - trial pair induces pairwise outcomes for all model pairs, so GG is complete and Cij=MNC_{ij}=MN for every i≠ji\ne j. Arena ranking removes this completeness assumption: only co-observed models contribute to an edge, and missing comparisons should not be interpreted as losses.

Which Ranking Families Transfer

Methods whose sufficient statistics are pairwise comparisons transfer most directly. For example, Bradley - Terry estimation on Arena data uses the same win counts as in the dense reduction. Let

pij=exp⁡(θi)exp⁡(θi)+exp⁡(θj),pji=1−pij.p_{ij}=\frac{\exp(\theta_i)}{\exp(\theta_i)+\exp(\theta_j)}, \qquad p_{ji}=1-p_{ij}.

The log-likelihood is

ℓBT(θ)=∑i<j(Wijlog⁡pij+Wjilog⁡pji).\ell_{\mathrm{BT}}(\theta) =\sum_{i<j}\bigl( W_{ij}\log p_{ij} +W_{ji}\log p_{ji} \bigr).

Tie-aware variants such as Davidson or Rao - Kupper additionally use TijT_{ij}. Sequential rating systems (Elo, Glicko, TrueSkill) can be run directly on the timestamped comparison stream, whereas the dense benchmark setting must first expand each question - trial slice into induced pairwise matches.

Graph and spectral methods also transfer after estimating pairwise preference probabilities on observed edges, for example

P^i≻j=Wij+12Tij+αCij+2α,{i,j}∈E,\widehat{P}_{i\succ j} =\frac{W_{ij}+\tfrac{1}{2}T_{ij}+\alpha}{C_{ij}+2\alpha}, \qquad \{i,j\}\in E,

where α≥0\alpha\ge0 is a smoothing constant. These edge weights can be used by PageRank, Rank Centrality, HodgeRank, α\alpha-Rank, and related graph-based procedures. Hodge-style decompositions are diagnostically useful because they separate a global ranking potential from cyclic residuals, which may reveal non-transitive preferences caused by prompt specialization, heterogeneous judges, or context-dependent model strengths. If an Arena compares more than two responses to the same prompt, the event may instead be represented by winner and loser sets (Ut,Vt)(U_t,V_t) and passed to listwise or setwise Luce-family models.

Pointwise metrics do not transfer without additional labels. Mean accuracy, Pass@kk, Bayes@NN, and standard IRT models require absolute model - item outcomes, whereas a preference log records only relative judgments. A response can win a comparison without being correct, or lose to a stronger response despite being acceptable. If the Arena protocol also records absolute signals - for example correctness, rubric scores, verifier labels, or categorical response features - then the problem becomes a masked benchmark rather than a pure preference arena. In that case, pointwise or IRT-style likelihoods can be written over the observed entries only,

Lmasked=∏Almn=1p(Rlmn∣θl,βm),\mathcal{L}_{\mathrm{masked}} =\prod_{A_{lmn}=1}p(R_{lmn}\mid \theta_l,\beta_m),

where AA is the observation mask. The mask is essential: zero-filling unobserved entries would conflate non-participation with failure and bias rankings toward frequently sampled models.

Sparsity also changes the role of regularization. In dense benchmarks, every model pair receives the same number of induced comparisons. In an Arena, low-degree models and disconnected components may be weakly identified or incomparable from data alone. Priors, anchor models, or a small dense benchmark pilot can therefore be more important for Arena ranking than for the controlled dense benchmark setting.

Evaluation Under Comparison Budgets

The stability protocol in section 3 extends to Arena logs by replacing the trial budget NN with a comparison budget. For a ranking method rr, let

πr(t)=r(D1:t),πr(T)=r(D1:T).\pi_r(t)=r(\mathcal{D}_{1:t}), \qquad \pi_r(T)=r(\mathcal{D}_{1:T}).

Agreement between πr(t)\pi_r(t) and πr(T)\pi_r(T), measured for example by Kendall's τb\tau_b, gives the Arena analogue of low-budget stability and convergence. Prefix evaluation preserves time order and measures how quickly a live leaderboard stabilizes; bootstrap evaluation resamples comparisons, prompts, or user sessions to quantify uncertainty. When observations share prompts, users, or judges, resampling at those higher levels is preferable to treating all comparisons as independent.

Arena rankings should be interpreted conditionally on the prompt distribution, judge population, model-selection policy, and decoding policy used to collect the log. Reporting should therefore include both rank estimates and diagnostics: connectivity and degree statistics of GarenaG_{\mathrm{arena}}, edge-count imbalance, posterior or bootstrap intervals for model strengths, and pairwise superiority probabilities such as Pr⁡(θi>θj∣Darena)\Pr(\theta_i>\theta_j\mid\mathcal{D}_{\mathrm{arena}}). Randomizing presentation order, maintaining anchor models, and stratifying by prompt domain make the estimand more transparent and reduce artifacts from side bias or adaptive sampling.

Arena ranking is therefore the sparse-comparison counterpart of the dense repeated-trial setting. The ranking estimators need not be redesigned when they operate on pairwise or setwise sufficient statistics; the observation model and uncertainty structure instead change because comparisons are sparse, non-uniform, and evolving. In Scorio, Arena logs correspond to sparse win/tie matrices or setwise events that can be processed by the paired-comparison, graph, or listwise rankers analyzed in this work.

Test-time scaling produces repeated stochastic outcomes per item, making LLM benchmarking closer to classical repeated-measurement settings than to single-run leaderboards. We summarize the main ranking families used in this work and their typical applications.

Paired-comparison and rating models.

Paired-comparison models represent comparisons through win/tie counts and infer latent strengths, with Bradley - Terry as a canonical likelihood-based model [10]. Practical systems often use online rating updates such as Elo and its extensions (e.g., Glicko) or fully Bayesian skill ratings such as TrueSkill [11, 12, 33]. For data with ties, common generalizations include Rao - Kupper and Davidson models [24, 25]. These models are widely used for preference aggregation in LLM leaderboards [7, 8], but are also natural in dense benchmarks once per-item outcomes are reduced to pairwise wins.

Listwise, setwise choice models.

When each trial yields an ordering over many items, listwise choice models such as Plackett - Luce provide a likelihood over permutations [34, 35]. Davidson - Luce extends setwise choice to allow ties within selected sets [36]. In our binary benchmark setting, each trial induces a two-level partition (solved vs. unsolved), so these models reduce to structured forms of pairwise likelihoods while still providing a principled view of aggregation.

IRT and difficulty-aware benchmarking.

Item response theory models couple model "ability" with item difficulty (and sometimes discrimination), with the Rasch and Birnbaum formulations as classic examples [18, 19]. IRT has recently been proposed as a way to disentangle model skill from benchmark composition in LLM evaluation [9]. When multiple trials per item are available, repeated-measures extensions and binomial-response formulations are natural [21, 22, 23], and difficulty reweighting has also been explored in NLP evaluation contexts [17].

Graph, spectral, and social-choice methods.

Beyond likelihood-based models, ranking from comparisons has a long tradition in social choice and graph-based aggregation. Voting rules such as Borda and Condorcet-style methods satisfy different axioms and can behave differently under noise and ties [13, 14, 15, 26]. Spectral and Markov-chain approaches derive scores from transition graphs, including PageRank and Rank Centrality [27, 28]; HodgeRank and related spectral methods interpret comparisons as edge flows and decompose them into global and cyclic components [29, 30]. AlphaRank was introduced for multi-agent evaluation with potentially non-transitive interactions [31], and related work studies open-ended evaluation dynamics [32]. We place these families in a common test-time-scaling benchmark setting and compare them under controlled increases in the number of repeated trials.

Experiment Setup and Reproducibility

Models and Datasets

Datasets.

We evaluate on four Olympiad-style math benchmarks: AIME'24 [48], AIME'25 [49], BrUMO'25 [50], and HMMT'25 [51]. For AIME'24 and AIME'25, we combine AIME I and AIME II from the corresponding year, yielding 3030 integer-answer problems per benchmark. For HMMT'25, we use the official February 2025 contest set, which spans algebra, geometry, number theory, and combinatorics. For BrUMO'25, we use the published 2025 problem sets from the tournament archive.

Models.

To reduce prompt-format confounds, we use provider-recommended chat templates (defaulting to DeepSeek/Qwen-style templates when no model-specific template is given) and shared decoding settings across models unless noted otherwise. We evaluate the 2020 models listed in table 22: Sky-T1-32B-Flash Sky-T1-32B-Flash [52] (Sky-T1 Flash release), Qwen Qwen3-30B-A3B-Thinking-2507 [53] (Qwen3 thinking model), DeepSeek DeepSeek-R1-Distill-Qwen-1.5B [54] (1.5B distilled reasoning model), gpt-oss gpt-oss-20b [55] (OpenAI open-weight model, evaluated with low, medium, and high Harmony reasoning effort under the default MXFP4 quantization), LIMO LIMO-v2 [56] (reasoning model), EXAONE EXAONE-4.0-1.2B [57] (hybrid reasoning/non-reasoning model), NVIDIA OpenReasoning-Nemotron-1.5B [58] (NVIDIA reasoning model), OpenThinker OpenThinker2-32B [59] and OpenThinker OpenThinker3-1.5B [59] (models trained from the OpenThoughts data recipes), Microsoft Phi-4-reasoning and Microsoft Phi-4-reasoning-plus [60], OpenR1 OpenR1-Distill-7B [61], FuseO1 FuseO1-DeepSeekR1-QwQ-SkyT1-Flash-32B-Preview [62], Light-R1 Light-R1-14B-DS [63], NVIDIA AceReason-Nemotron-1.1-7B [64], NVIDIA NVIDIA-Nemotron-Nano-9B-v2 [65], Qwen Qwen3-4B-Thinking-2507 [53], and Bespoke-Stratos Bespoke-Stratos-7B [66].

Table 22. Mapping between model IDs, full model names, and the shortened names used in figures and legends.
IDModelShort name
1DeepSeek DeepSeek-R1-Distill-Qwen-1.5BDS-R1-Qwen
2LIMO LIMO-v2LIMO-v2
3OpenThinker OpenThinker2-32BOpenThinker2
4OpenThinker OpenThinker3-1.5BOpenThinker3
5Qwen Qwen3-30B-A3B-Thinking-2507Qwen3-Thinking
6Sky-T1-32B-Flash Sky-T1-32B-FlashSky-T1-Flash
7gpt-oss gpt-oss-20b_highgpt-oss-high
8gpt-oss gpt-oss-20b_lowgpt-oss-low
9gpt-oss gpt-oss-20b_mediumgpt-oss-medium
10EXAONE EXAONE-4.0-1.2BEXAONE-4.0
11NVIDIA OpenReasoning-Nemotron-1.5BOR-Nemotron
12Microsoft Phi-4-reasoningPhi-4
13Microsoft Phi-4-reasoning-plusPhi-4-plus
14OpenR1 OpenR1-Distill-7BOR1-Distill
15FuseO1 FuseO1-DeepSeekR1-QwQ- SkyT1-Flash-32B-PreviewFuseO1-DS-QwQ-SkyT1
16Light-R1 Light-R1-14B-DSLight-R1-DS
17NVIDIA AceReason-Nemotron-1.1-7BAR-Nemotron
18NVIDIA NVIDIA-Nemotron-Nano-9B-v2NVIDIA-Nemotron
19Qwen Qwen3-4B-Thinking-2507Qwen3-4B
20Bespoke-Stratos Bespoke-Stratos-7BBespoke
Prompting.

We use provider-recommended prompt templates for each model. For most models, we adopt the standard DeepSeek/Qwen-style prompt, "Please reason step by step, and put your final answer within boxed{}." For gpt-oss gpt-oss-20b, we use the OpenAI Harmony prompt template, which specifies three discrete levels of reasoning effort. For NVIDIA OpenReasoning-Nemotron-1.5B, we use the task-specific prompt, "Solve the following math problem. Make sure to put the answer (and only the answer) inside boxed{}."

Reproducibility

For stochastic runs, we use top-pp sampling with temperature 0.60.6, p=0.95p=0.95, batch size 11, and random seeds 12341234 through 13131313, yielding N=80N=80 trials per dataset - model pair. All models are served with vLLM (PagedAttention) [67] in bf16 precision, except releases that require MXFP4 quantization (e.g., gpt-oss). We record log-probabilities for both input prompts and generated tokens, with max_tokens set to 32,76832{,}768. All experiments run on clusters equipped with 8×8\times NVIDIA H200 GPUs (141GB per GPU).

Computational Cost and Token Statistics

We evaluate 20 models across four benchmarks, with 80 trials per model and 30 questions per benchmark, for a total of 192,000 independent inference runs. The full evaluation requires 7,445 GPU-hours (approximately 310 GPU-days) and generates 2.96B tokens (2,963,318,176 total); table 23 reports the task-level totals. Of these tokens, 37M (1.2%) are prompt tokens and 2.93B (98.8%) are completion tokens, for an average of 15,434 tokens per query. Among the four benchmarks, HMMT'25 is the most computationally expensive at 2,217 GPU-hours, whereas BrUMO'25 is the least expensive at 1,651 GPU-hours. Across model configurations, gpt-oss-20b-low is the most efficient (48.4 GPU-hours for 9,600 queries) and LIMO-v2 the least efficient (894.3 GPU-hours for the same workload), with a corpus-wide average of 139.6 seconds per query.

Table 23. Task-level computational cost aggregated over 20 models, 80 trials, four tasks, and 30 questions per task. Token counts correspond to completion tokens only.
TaskInference Time (hours)Completion Tokens (M)
AIME'241,699.4680.0
AIME'251,878.4728.3
HMMT'252,216.5851.2
BrUMO'251,650.9666.9
TOTAL7,445.22,926.4

Rank Correlation Metrics

Kendall's tau

Kendall's tau (τ\tau) [68] measures ordinal agreement between two rankings through pairwise concordance and discordance. For rankings of nn items, let ncn_c and ndn_d denote the numbers of concordant and discordant pairs, let n0=n(n−1)/2n_0 = n(n-1)/2 be the total number of pairs, and let n1n_1 and n2n_2 be the numbers of tied pairs in the two rankings. The two common variants are

Tau-a:τa=nc−ndn0,\begin{aligned}\text{Tau-a:}\quad \tau_a &= \frac{n_c - n_d}{n_0},\end{aligned} (18)
Tau-b:τb=nc−nd(n0−n1)(n0−n2).\begin{aligned}\text{Tau-b:}\quad \tau_b &= \frac{n_c - n_d}{\sqrt{(n_0 - n_1)(n_0 - n_2)}}.\end{aligned} (19)

Tau-a ignores ties, whereas Tau-b corrects for them. Because ties are common in our setting, we use τb\tau_b throughout.

Scorio, Open-Source Library for LLM Ranking

Scorio is a Python library for ranking LLMs from repeated-trial benchmark evaluations under test-time scaling. It provides a unified interface for mapping the response tensor R∈{0,1}L×M×N\mathbf{R}\in\{0,1\}^{L\times M\times N} (and, where relevant, optional prior outcomes) to model scores and rankings across evaluation metrics, probabilistic paired-comparison and rating systems, voting rules, listwise choice models, item response theory, and graph- or spectral-based methods. The library is distributed through PyPI as scorio.

All ranking methods in Scorio operate on the response tensor R\mathbf{R}, where LL is the number of models, MM the number of questions, and NN the number of trials per question. The implementation represents this tensor as a NumPy array of shape (L, M, N). Listing 1 gives a minimal example of constructing R\mathbf{R} and calling a basic ranking method.

Listing 1. Constructing the response tensor and computing rankings with Scorio.
import numpy as np
from scorio import rank

# Binary response tensor: L=3 models, M=4 questions, N=5 trials
R = np.random.randint(0, 2, size=(3, 4, 5))

# Rank by mean accuracy
rankings = rank.avg(R)

# Return both rankings and scores
rankings, scores = rank.avg(R, return_scores=True)

The rank module uses a common interface: each function takes the tensor R\mathbf{R} as the first argument, returns a ranking array of shape (L,), and accepts an optional return_scores=True flag to additionally return the underlying scores. Rankings are 11-indexed, with lower values indicating better models.

Scorio implements ranking methods from several families. Listing 2 illustrates evaluation-based methods, including the Pass@kk family that quantifies how reliably models solve questions within kk sampled trials.

Listing 2. Evaluation-based ranking methods.
# Pass@k: probability at least 1 of k draws succeeds
rankings, scores = rank.pass_at_k(R, k=3, return_scores=True)

# G-Pass@k with threshold tau
rankings = rank.g_pass_at_k_tau(R, k=5, tau=0.6)

# Bayesian posterior ranking with optional prior outcomes
R0 = np.random.randint(0, 2, size=(3, 4, 2))  # prior data
rankings = rank.bayes(R, R0=R0)

The bayes method generalizes beyond binary correctness to categorical outcomes Rlmn∈{0,…,C}R_{lmn}\in\{0,\dots,C\} via a weight vector w∈RC+1\mathbf{w}\in\mathbb{R}^{C+1} that maps each category to a score. It also accepts an optional prior tensor R0\mathbf{R}_0 that incorporates outcomes from a different evaluation setting (e.g., greedy decoding) as a Bayesian prior. Listing 3 gives examples of both cases.

Listing 3. Bayes@NN with categorical outcomes and greedy prior.
# Categorical outcomes: 0=wrong, 1=partial, 2=correct
# L=3 models, M=4 questions, N=5 trials
R_cat = np.random.randint(0, 3, size=(3, 4, 5))

# Weight vector mapping categories to scores
w = np.array([0.0, 0.5, 1.0])

rankings, scores = rank.bayes(R_cat, w=w,
                              return_scores=True)

# Using greedy decoding results as Bayesian prior
# R0 shape (M, D): shared prior across all models
R0_greedy = np.random.randint(0, 3, size=(4, 2))
rankings = rank.bayes(R_cat, w=w, R0=R0_greedy)

# Conservative ranking via posterior quantile
rankings = rank.bayes(R_cat, w=w, R0=R0_greedy,
                      quantile=0.05)

For probabilistic paired-comparison models, Scorio implements the Bradley - Terry model and its extensions, as well as Elo and TrueSkill rating systems (listing 4). These methods construct pairwise comparisons from R\mathbf{R} and estimate latent strength parameters.

Listing 4. Paired-comparison and rating system methods.
# Bradley-Terry maximum likelihood
rankings, scores = rank.bradley_terry(R, return_scores=True)

# Bradley-Terry with MAP regularization
rankings = rank.bradley_terry_map(R, prior=1.0)

# Elo rating system
rankings, scores = rank.elo(R, K=32.0, return_scores=True)

# TrueSkill Bayesian rating
rankings = rank.trueskill(R)

Graph-based and spectral methods rank models by analyzing the structure of a pairwise comparison graph derived from R\mathbf{R}, as shown in listing 5.

Listing 5. Graph-based ranking methods.
# PageRank on the pairwise win-probability graph
rankings, scores = rank.pagerank(R, damping=0.85,
                                 return_scores=True)

# Spectral ranking (principal eigenvector)
rankings = rank.spectral(R)

# Rank centrality via Markov chain stationary distribution
rankings = rank.rank_centrality(R)

A family-wise list of ranking methods is given in appendix J.1, and the exact method configurations used in our experiments are reported in appendix J.2.

Ranking Methods

Pointwise Methods

Mean accuracy.

The simplest pointwise score is the mean accuracy

slmean:=1M∑m=1Mp^lm,s_l^{\mathrm{mean}} := \frac{1}{M}\sum_{m=1}^M \widehat{p}_{lm}, (20)

which corresponds to avg in Scorio.

Inverse-difficulty weighting.

To emphasize hard questions, inverse_difficulty weights each question by the inverse of its global solve rate pm:=1LN∑l,nRlmnp_m := \frac{1}{LN}\sum_{l,n} R_{lmn}:

wm∝1clip⁡(pm,ϵ,1−ϵ),slinv-diff:=∑m=1Mwm p^lm,\begin{aligned} w_m &\propto \frac{1}{\operatorname{clip}(p_m,\epsilon,1-\epsilon)},\\ s_l^{\mathrm{inv\text{-}diff}} &:= \sum_{m=1}^M w_m\,\widehat{p}_{lm}, \end{aligned} (21)

with weights normalized to ∑mwm=1\sum_m w_m = 1.

Algorithm 1. Pointwise scoring (mean and inverse-difficulty)
Require R∈{0,1}L×M×NR\in\{0,1\}^{L\times M\times N}, ϵ>0\epsilon>0
Ensure Scores s∈RLs\in\mathbb{R}^L
Compute p^lm←1N∑n=1NRlmn\widehat{p}_{lm}\gets \frac{1}{N}\sum_{n=1}^N R_{lmn}
Mean: sl←1M∑m=1Mp^lms_l \gets \frac{1}{M}\sum_{m=1}^M \widehat{p}_{lm}
Inv-diff: compute pm←1LN∑l,nRlmnp_m\gets \frac{1}{LN}\sum_{l,n}R_{lmn}
Set wm∝1/clip⁡(pm,ϵ,1−ϵ)w_m\propto 1/\operatorname{clip}(p_m,\epsilon,1-\epsilon) and normalize to ∑mwm=1\sum_m w_m=1
sl←∑mwm p^lms_l \gets \sum_{m} w_m\,\widehat{p}_{lm}

Evaluation-metric Methods

These methods rank models by evaluation metrics computed from per-question trial outcomes. The simplest baseline is mean accuracy (avg; appendix J.1.1); we next define Pass@kk-family metrics and Bayes@NN. For a fixed model ll, define the per-question success counts νlm:=∑n=1NRlmn\nu_{lm}:=\sum_{n=1}^N R_{lmn}. Each metric defines a per-question score f(νlm;N)f(\nu_{lm};N) (or f(νlm;N,k,τ)f(\nu_{lm};N,k,\tau)) and then averages across questions.

Pass@kk (pass_at_k).

Pass@kk [1] is the probability that at least one of kk samples is correct. For each question mm,

Pass@klm:=1−(N−νlmk)(Nk),\mathrm{Pass@}k_{lm} := 1 - \frac{\binom{N-\nu_{lm}}{k}}{\binom{N}{k}}, (22)

and the model-level score is slPass@k:=1M∑m=1MPass@klms_l^{\mathrm{Pass@}k} := \frac{1}{M}\sum_{m=1}^M \mathrm{Pass@}k_{lm}.

Pass-hat@k / G-Pass@k (pass_hat_k).

This metric (also called G-Pass@k in parts of the recent LLM evaluation literature [69]) is the probability that all kk selected samples are correct:

Pass@k^lm:=(νlmk)(Nk),\widehat{\mathrm{Pass@}k}_{lm} := \frac{\binom{\nu_{lm}}{k}}{\binom{N}{k}}, (23)

with slPass@k^:=1M∑m=1MPass@k^lms_l^{\widehat{\mathrm{Pass@}k}} := \frac{1}{M}\sum_{m=1}^M \widehat{\mathrm{Pass@}k}_{lm}.

G-Pass@kτ_\tau (g_pass_at_k_tau).

G-Pass@kτ_\tau [42] generalizes these metrics by requiring at least j0:=⌈τk⌉j_0:=\lceil\tau k\rceil successes among the kk selected samples. Let Xlm∼Hypergeom(N,νlm,k)X_{lm}\sim\mathrm{Hypergeom}(N,\nu_{lm},k) be the number of successes in a draw of size kk without replacement; then

G-Pass@kτ,lm:=Pr⁡(Xlm≥j0)=∑j=j0k(νlmj)(N−νlmk−j)(Nk),\begin{aligned} \mathrm{G\text{-}Pass@}k_{\tau,lm} &:= \Pr(X_{lm}\ge j_0)\\ &= \sum_{j=j_0}^{k}\frac{\binom{\nu_{lm}}{j}\binom{N-\nu_{lm}}{k-j}}{\binom{N}{k}}, \end{aligned} (24)

and slG-Pass@kτ:=1M∑m=1MG-Pass@kτ,lms_l^{\mathrm{G\text{-}Pass@}k_\tau}:=\frac{1}{M}\sum_{m=1}^M \mathrm{G\text{-}Pass@}k_{\tau,lm}. Scorio defines the endpoint τ=0\tau=0 to recover Pass@kk (and for any τ∈(0,1/k]\tau\in(0,1/k] the threshold j0=⌈τk⌉j_0=\lceil\tau k\rceil equals 11, so the expression matches Pass@kk), while τ=1\tau=1 recovers Pass-hat@k.

mG-Pass@k (mg_pass_at_k).

mG-Pass@k [42] aggregates G-Pass@kτ_\tau over τ∈[0.5,1]\tau\in[0.5,1]. In Scorio, we use the equivalent expectation form

mG-Pass@klm:=2k E ⁣[(Xlm−m0)+],m0:=⌈k2⌉,\begin{aligned} \mathrm{mG\text{-}Pass@}k_{lm} &:= \frac{2}{k}\,\mathbb{E}\!\left[(X_{lm}-m_0)_+\right],\\ m_0 &:=\left\lceil \tfrac{k}{2}\right\rceil, \end{aligned} (25)

where (x)+:=max⁡(x,0)(x)_+:=\max(x,0) and Xlm∼Hypergeom(N,νlm,k)X_{lm}\sim\mathrm{Hypergeom}(N,\nu_{lm},k). The model-level score is slmG-Pass@k:=1M∑m=1MmG-Pass@klms_l^{\mathrm{mG\text{-}Pass@}k}:=\frac{1}{M}\sum_{m=1}^M \mathrm{mG\text{-}Pass@}k_{lm}.

Bayes@NN (bayes).

Bayes@NN [16] applies to multi-category outcomes Rlmn∈{0,…,C}R_{lmn}\in\{0,\dots,C\} with a weight vector w∈RC+1w\in\mathbb{R}^{C+1}. For a fixed model ll and question mm, let nmk:=∑n=1N1{Rlmn=k}n_{mk}:=\sum_{n=1}^N \mathbf{1}\{R_{lmn}=k\} be category counts. Optionally, a prior outcome matrix R0∈{0,…,C}M×DR_0\in\{0,\dots,C\}^{M\times D} contributes pseudo-counts nmk0:=1+∑d=1D1{(R0)md=k}n^0_{mk}:=1+\sum_{d=1}^D \mathbf{1}\{(R_0)_{md}=k\} (a Dirichlet(1,…,1)(1,\dots,1) prior), giving νmk:=nmk+nmk0\nu_{mk}:=n_{mk}+n^0_{mk} and T:=1+C+D+NT:=1+C+D+N. Bayes@NN returns a posterior mean μl\mu_l and uncertainty σl\sigma_l of the weighted score:

μl=w0+1M T∑m=1M∑k=0Cνmk(wk−w0),\mu_l = w_0 + \frac{1}{M\,T}\sum_{m=1}^M\sum_{k=0}^{C}\nu_{mk}(w_k-w_0), (26)
σl=(1M2(T+1)∑m=1M[∑kνmkT(wk−w0)2−(∑kνmkT(wk−w0))2])1/2.\begin{aligned} \sigma_l &= \Bigl(\frac{1}{M^2(T+1)}\sum_{m=1}^M\Bigl[ \sum_k \frac{\nu_{mk}}{T}(w_k-w_0)^2\\ &\qquad\qquad -\Bigl(\sum_k \frac{\nu_{mk}}{T}(w_k-w_0)\Bigr)^2 \Bigr]\Bigr)^{1/2}. \end{aligned} (27)

Scorio ranks by μl\mu_l (default) or by a conservative normal-quantile score μl+Φ−1(q)σl\mu_l+\Phi^{-1}(q)\sigma_l for a chosen q∈[0,1]q\in[0,1].

Bayesian Methods

Thompson sampling ranking (thompson).

Thompson sampling [70, 71] ranks by Monte Carlo samples from a conjugate Beta - Binomial posterior over each model's aggregate success probability. We model pl∼Beta(α,β)p_l\sim\mathrm{Beta}(\alpha,\beta) and treat all MNMN trials as i.i.d. Bernoulli outcomes [38]. Let Sl:=∑m=1M∑n=1NRlmnS_l:=\sum_{m=1}^M\sum_{n=1}^N R_{lmn} be the total number of successes for model ll; then

pl∣R∼Beta ⁣(α+Sl,  β+MN−Sl).p_l \mid \mathbf{R} \sim \mathrm{Beta}\!\left(\alpha + S_l,\;\beta + MN - S_l\right). (28)

For t=1,…,Tt=1,\dots,T we draw pl(t)∼pl∣Rp_l^{(t)}\sim p_l\mid\mathbf{R} independently for each model, compute the induced rank rl(t)∈{1,…,L}r_l^{(t)}\in\{1,\dots,L\} (smaller is better), and score by the negative average rank

slTS:=−1T∑t=1Trl(t).s_l^{\mathrm{TS}} := -\frac{1}{T}\sum_{t=1}^T r_l^{(t)}. (29)
Bayesian Bradley - Terry via MCMC (bayesian_mcmc).

To obtain a full Bayesian posterior over paired-comparison strengths, we combine the Bradley - Terry likelihood [10] with a Gaussian prior and approximate the posterior with Metropolis - Hastings sampling [72, 73]. We first form decisive win counts

Wij:=∑m=1M∑n=1N1{Rimn=1,  Rjmn=0},W_{ij} := \sum_{m=1}^M\sum_{n=1}^N \mathbf{1}\{R_{imn}=1,\;R_{jmn}=0\}, (30)

ignoring ties (both correct or both incorrect). Parameterizing πi=exp⁡(θi)\pi_i=\exp(\theta_i), the BT likelihood is

Pr⁡(i≻j∣θ)=exp⁡(θi)exp⁡(θi)+exp⁡(θj),log⁡p(W∣θ)=∑i≠jWijlog⁡Pr⁡(i≻j∣θ),\begin{aligned} \Pr(i\succ j \mid \theta) &= \frac{\exp(\theta_i)}{\exp(\theta_i)+\exp(\theta_j)},\\ \log p(\mathbf{W}\mid \theta) &= \sum_{i\neq j} W_{ij}\log \Pr(i\succ j \mid \theta), \end{aligned} (31)

with an independent prior θi∼N(0,σ2)\theta_i\sim\mathcal{N}(0,\sigma^2) [39]. We sample from p(θ∣W)p(\theta\mid \mathbf{W}) and rank models by the posterior mean score siMCMC:=E[θi∣W]s_i^{\mathrm{MCMC}}:=\mathbb{E}[\theta_i\mid \mathbf{W}].

Voting-based Methods

Voting rules aggregate per-question preferences into a global ranking. To adapt them to our test-time-scaling setting, we treat each question mm as a "voter" that ranks models by their per-question solve frequency across trials:

klm:=∑n=1NRlmn∈{0,1,…,N}.k_{lm} := \sum_{n=1}^N R_{lmn}\in\{0,1,\dots,N\}. (32)

When N=1N=1, each question induces only a two-level ranking (correct vs. incorrect), so Borda/Copeland reduce to (ties of) accuracy-based ordering; when N>1N>1 these rules exploit the additional resolution from klmk_{lm}.

Borda count.

For each question mm, let rlm∈{1,…,L}r_{lm}\in\{1,\dots,L\} be the (tie-averaged) rank of model ll when sorting k⋅mk_{\cdot m} in descending order (smaller rank is better). The Borda score is

slBorda:=∑m=1M(L−rlm),s_l^{\mathrm{Borda}} := \sum_{m=1}^M (L - r_{lm}), (33)

which assigns (L−1)(L-1) points for a unique first place and 00 for a unique last place, with ties receiving the average of the tied positions [13, 26].

Copeland.

For each pair (i,j)(i,j), define the number of questions that prefer ii to jj as Wij(q):=∑mI[kim>kjm]W^{(q)}_{ij}:=\sum_m \mathbb{I}[k_{im}>k_{jm}]. Copeland declares ii to beat jj if Wij(q)>Wji(q)W^{(q)}_{ij}>W^{(q)}_{ji} and scores each model by net pairwise dominance:

siCopeland:=∑j≠isign ⁣(Wij(q)−Wji(q)),s_i^{\mathrm{Copeland}} := \sum_{j\neq i}\mathrm{sign}\!\left(W^{(q)}_{ij}-W^{(q)}_{ji}\right), (34)

where sign(0)=0\mathrm{sign}(0)=0 [74, 26].

Win rate.

Using the same question-level win counts W(q)W^{(q)}, define a model's win rate as the fraction of decisive pairwise outcomes it wins:

siwinrate:=∑j≠iWij(q)∑j≠i(Wij(q)+Wji(q)),s_i^{\mathrm{winrate}} := \frac{\sum_{j\neq i} W^{(q)}_{ij}}{\sum_{j\neq i} \left(W^{(q)}_{ij}+W^{(q)}_{ji}\right)}, (35)

with the convention siwinrate=0.5s_i^{\mathrm{winrate}}=0.5 if the denominator is zero.

Condorcet-style pairwise-majority rules.

Many voting rules are defined from an aggregated pairwise preference matrix. To incorporate per-question ties when kim=kjmk_{im}=k_{jm}, we define

Pij(q):=∑m=1M(I[kim>kjm]+12I[kim=kjm]),P^{(q)}_{ij} := \sum_{m=1}^M \Big(\mathbb{I}[k_{im}>k_{jm}] + \tfrac{1}{2}\mathbb{I}[k_{im}=k_{jm}]\Big), (36)

so that Pij(q)+Pji(q)=MP^{(q)}_{ij}+P^{(q)}_{ji}=M. Let margins be Δij:=Pij(q)−Pji(q)\Delta_{ij}:=P^{(q)}_{ij}-P^{(q)}_{ji}.

Minimax (Simpson - Kramer).

The minimax score is based on a model's worst pairwise defeat:

siminimax:=−max⁡j≠imax⁡(0,Δji),s_i^{\mathrm{minimax}} := -\max_{j\neq i}\max(0,\Delta_{ji}), (37)

and ranks models by the size of their worst defeat (closer to 00 is better) [26].

Schulze (beatpath).

Schulze computes strongest-path strengths pijp_{ij} in the directed graph of pairwise victories and ranks ii above jj if pij>pjip_{ij}>p_{ji} [75, 26].

Ranked Pairs (Tideman).

Ranked Pairs sorts pairwise victories by strength (e.g., margin Δij\Delta_{ij}), then locks them in that order whenever doing so does not introduce a cycle; the resulting acyclic dominance graph induces a ranking [76, 26].

Kemeny - Young.

Kemeny - Young returns an ordering π\pi that maximizes agreement with the pairwise preferences:

π∈arg⁡max⁡total orders π  ∑i≺πjPij(q),\pi \in \arg\max_{\text{total orders } \pi}\;\sum_{i\prec_\pi j} P^{(q)}_{ij}, (38)

which is equivalent to a maximum-likelihood ranking under certain noise models and is a classic Condorcet extension [77, 78, 26]. (Exact optimization is NP-hard in general; we solve the induced linear ordering problem via MILP for the problem sizes in this paper.)

Borda elimination rules (Nanson and Baldwin).

Nanson's method iteratively recomputes Borda scores over remaining candidates and removes those below the mean, while Baldwin's method removes the lowest Borda scorer(s) each round [79, 80, 26].

Majority Judgment.

Majority Judgment treats klm∈{0,…,N}k_{lm}\in\{0,\dots,N\} as discrete grades and ranks models by their median grade, breaking ties using the majority-gauge rule [81].

Algorithm 2. Voting rules on per-question trial counts
Require R∈{0,1}L×M×NR\in\{0,1\}^{L\times M\times N}
Ensure Borda scores sBordas^{\mathrm{Borda}}, Copeland scores sCopelands^{\mathrm{Copeland}}, win-rate scores swinrates^{\mathrm{winrate}}
Compute klm←∑n=1NRlmnk_{lm}\gets \sum_{n=1}^N R_{lmn}
sBorda←0s^{\mathrm{Borda}}\gets 0; sCopeland←0s^{\mathrm{Copeland}}\gets 0; initialize W(q)←0W^{(q)}\gets 0
For m=1m=1 to MM
Rank models by k⋅mk_{\cdot m} (descending) with average-tie ranks r⋅mr_{\cdot m}
slBorda+=L−rlms_l^{\mathrm{Borda}} \mathrel{+}= L - r_{lm} for all ll
EndFor
For 1≤i<j≤L1\le i<j\le L
Wij(q)←∑mI[kim>kjm]W^{(q)}_{ij}\gets \sum_m \mathbb{I}[k_{im}>k_{jm}]
Wji(q)←∑mI[kjm>kim]W^{(q)}_{ji}\gets \sum_m \mathbb{I}[k_{jm}>k_{im}]
If Wij(q)>Wji(q)W^{(q)}_{ij}>W^{(q)}_{ji} siCopeland+=1s_i^{\mathrm{Copeland}}\mathrel{+}=1; sjCopeland−=1s_j^{\mathrm{Copeland}}\mathrel{-}=1
ElsIf Wji(q)>Wij(q)W^{(q)}_{ji}>W^{(q)}_{ij} siCopeland−=1s_i^{\mathrm{Copeland}}\mathrel{-}=1; sjCopeland+=1s_j^{\mathrm{Copeland}}\mathrel{+}=1
EndIf
EndFor
siwinrate←∑j≠iWij(q)∑j≠i(Wij(q)+Wji(q))s_i^{\mathrm{winrate}}\gets \frac{\sum_{j\neq i} W^{(q)}_{ij}}{\sum_{j\neq i}(W^{(q)}_{ij}+W^{(q)}_{ji})} (or 0.50.5 if denominator is 00)

Paired-comparison Probabilistic Models

These methods first reduce R\mathbf{R} to pairwise win/tie counts between models, then fit a parametric paired-comparison model. For each ordered pair (i,j)(i,j), define wins WijW_{ij} and ties TijT_{ij} as in section 2.2 (pairwise representation).

Bradley - Terry (BT).

The BT model [10] assigns each model a positive strength πi>0\pi_i>0 and assumes

Pr⁡(i≻j)=πiπi+πj.\Pr(i\succ j) = \frac{\pi_i}{\pi_i+\pi_j}. (39)

Given win counts WijW_{ij}, the log-likelihood is

log⁡p(W∣π)=∑i≠jWij[log⁡πi−log⁡(πi+πj)],\log p(\mathbf{W}\mid \pi) = \sum_{i\neq j} W_{ij}\Big[\log \pi_i - \log(\pi_i+\pi_j)\Big], (40)

with identifiability enforced by centering log-strengths. Scorio provides ML (bradley_terry) and MAP (bradley_terry_map) estimation; MAP adds a prior penalty on log-strengths (e.g., Gaussian) [39].

Tie extensions.

In our binary setting, a pairwise tie occurs when both models are correct or both are incorrect on the same question - trial. Scorio implements two classic tie models:

  • Davidson [25]: adds a tie parameter and models (i≻j)(i\succ j), (j≻i)(j\succ i), and (i∼j)(i\sim j) explicitly (bradley_terry_davidson, bradley_terry_davidson_map).

  • Rao - Kupper [24]: alternative tie parameterization via κ≥1\kappa\ge 1 (rao_kupper, rao_kupper_map).

For Davidson, with tie parameter ν>0\nu>0,

Pr⁡(i≻j)=πiπi+πj+νπiπj,Pr⁡(j≻i)=πjπi+πj+νπiπj,Pr⁡(i∼j)=νπiπjπi+πj+νπiπj.\begin{aligned} \Pr(i\succ j) &= \frac{\pi_i}{\pi_i+\pi_j+\nu\sqrt{\pi_i\pi_j}},\\ \Pr(j\succ i) &= \frac{\pi_j}{\pi_i+\pi_j+\nu\sqrt{\pi_i\pi_j}},\\ \Pr(i\sim j) &= \frac{\nu\sqrt{\pi_i\pi_j}}{\pi_i+\pi_j+\nu\sqrt{\pi_i\pi_j}}. \end{aligned} (41)

For Rao - Kupper, with κ≥1\kappa\ge 1,

Pr⁡(i≻j)=πiπi+κπj,Pr⁡(j≻i)=πjκπi+πj,Pr⁡(i∼j)=(κ2−1)πiπj(πi+κπj)(κπi+πj).\begin{aligned} \Pr(i\succ j) &= \frac{\pi_i}{\pi_i+\kappa\pi_j},\\ \Pr(j\succ i) &= \frac{\pi_j}{\kappa\pi_i+\pi_j},\\ \Pr(i\sim j) &= \frac{(\kappa^2-1)\pi_i\pi_j}{(\pi_i+\kappa\pi_j)(\kappa\pi_i+\pi_j)}. \end{aligned} (42)
Algorithm 3. Paired-comparison models (BT, Davidson, Rao - Kupper) via ML/MAP
Require R∈{0,1}L×M×NR\in\{0,1\}^{L\times M\times N}; model family; optional prior penalty on log-strengths; max iterations TT
Ensure Scores (strengths) π^∈R+L\hat{\pi}\in\mathbb{R}_+^L
Compute pairwise win/tie counts (Wij,Tij)(W_{ij},T_{ij}) from RR
Parameterize strengths by log-strengths θi=log⁡πi\theta_i=\log \pi_i and enforce identifiability by centering: θ←θ−1L∑iθi\theta\leftarrow \theta-\frac{1}{L}\sum_i\theta_i
Define the family-specific log-likelihood log⁡p(W,T∣θ,tie-params)\log p(W,T\mid \theta,\text{tie-params})
Define objective L=−log⁡p(⋅)+prior(θ)\mathcal{L}=-\log p(\cdot)+\text{prior}(\theta) (prior term is 00 for ML)
Optimize L\mathcal{L} with L-BFGS for up to TT iterations
Return π^i=exp⁡(θ^i)\hat{\pi}_i=\exp(\hat{\theta}_i) as scores (larger is better)

Sequential Rating Systems

Sequential rating systems process a stream of head-to-head "matches" rather than aggregating all pairwise outcomes into a single count matrix. In our benchmark setting, the natural match stream is induced by each question - trial (m,n)(m,n): for every pair of models (i,j)(i,j), we observe a binary outcome pair (Rimn,Rjmn)∈{0,1}2(R_{imn},R_{jmn})\in\{0,1\}^2 and declare ii to beat jj if (1,0)(1,0) and jj to beat ii if (0,1)(0,1). When (Rimn,Rjmn)∈{(1,1),(0,0)}(R_{imn},R_{jmn})\in\{(1,1),(0,0)\}, the comparison is a tie; Scorio exposes tie-handling policies (e.g., treat ties as draws or ignore certain ties) for these methods.

Elo.

Elo [11] maintains a scalar rating rir_i for each model. For a match between ii and jj, define the expected score

Eij:=11+10(rj−ri)/400,E_{ij} := \frac{1}{1 + 10^{(r_j-r_i)/400}}, (43)

and let Sij∈{0,12,1}S_{ij}\in\{0,\tfrac{1}{2},1\} be the realized match score for ii against jj (win/draw/loss, depending on the tie-handling rule). The sequential Elo update is

ri←ri+K(Sij−Eij),rj←rj+K((1−Sij)−(1−Eij)),\begin{aligned} r_i &\leftarrow r_i + K(S_{ij}-E_{ij}),\\ r_j &\leftarrow r_j + K((1-S_{ij})-(1-E_{ij})), \end{aligned} (44)

with learning rate K>0K>0 (elo in Scorio). Because the updates are sequential, the final ratings can depend on the order in which the match stream is processed.

Glicko.

Glicko [12] augments Elo with an uncertainty parameter (rating deviation) RDi\mathrm{RD}_i and updates ratings using batches of matches within rating periods. In our implementation, each question - trial (m,n)(m,n) constitutes one rating period containing all pairwise matches on that (m,n)(m,n). Define q:=ln⁡(10)/400q:=\ln(10)/400 and

g(RD):=11+3q2RD2π2.g(\mathrm{RD}) := \frac{1}{\sqrt{1+\frac{3q^2\mathrm{RD}^2}{\pi^2}}}. (45)

For a player ii in a rating period with opponents j∈Oij\in\mathcal{O}_i and outcomes SijS_{ij}, define expected scores

Eij:=11+10−g(RDj)(ri−rj)/400,E_{ij} := \frac{1}{1 + 10^{-g(\mathrm{RD}_j)(r_i-r_j)/400}}, (46)

and

di2:=(q2∑j∈Oig(RDj)2Eij(1−Eij))−1.d_i^2 := \left(q^2\sum_{j\in\mathcal{O}_i} g(\mathrm{RD}_j)^2 E_{ij}(1-E_{ij})\right)^{-1}. (47)

The Glicko updates are

RDi′:=(1RDi2+1di2)−1/2,ri′:=ri+q1RDi2+1di2∑j∈Oig(RDj)(Sij−Eij),\begin{aligned} \mathrm{RD}_i' &:= \left(\frac{1}{\mathrm{RD}_i^2} + \frac{1}{d_i^2}\right)^{-1/2},\\ r_i' &:= r_i + \frac{q}{\frac{1}{\mathrm{RD}_i^2}+\frac{1}{d_i^2}} \sum_{j\in\mathcal{O}_i} g(\mathrm{RD}_j)(S_{ij}-E_{ij}), \end{aligned} (48)

with optional RD\mathrm{RD} inflation between rating periods and a maximum RD\mathrm{RD} cap (as in the original Glicko specification). This corresponds to glicko in Scorio; we rank by ri′r_i' (larger is better), and RDi′\mathrm{RD}_i' can be used as an uncertainty summary.

TrueSkill.

TrueSkill [33] is a Bayesian rating system that models each model's latent skill as a Gaussian N(μi,σi2)\mathcal{N}(\mu_i,\sigma_i^2) and updates (μi,σi)(\mu_i,\sigma_i) after each match using approximate inference. In Scorio, we apply a two-player TrueSkill update to each decisive (1,0)(1,0) or (0,1)(0,1) pairwise match in the induced stream (ties are ignored) and return the final μi\mu_i as the score (trueskill); a per-round dynamics parameter τ\tau inflates σ\sigma between rounds to model drift.

Listwise / Setwise Choice Models (Luce Family)

Unlike pairwise models, these methods operate on setwise events induced by each question - trial (m,n)(m,n). Define the winner and loser sets

Umn:={l:Rlmn=1},Vmn:={l:Rlmn=0}.\begin{aligned} U_{mn} &:= \{l: R_{lmn}=1\},\\ V_{mn} &:= \{l: R_{lmn}=0\}. \end{aligned} (49)

If Umn=∅U_{mn}=\emptyset or Umn=LU_{mn}=\mathcal{L}, the event contains no ranking information and is discarded.

Plackett - Luce (PL).

The PL model [34, 35] is a listwise generalization of BT for full rankings. In our binary setting we apply PL to the pairwise win matrix (equivalently BT) and estimate strengths using the MM update from [37] (plackett_luce, plackett_luce_map).

Davidson - Luce (setwise ties).

Davidson - Luce [36] models the probability of a tied winner set UmnU_{mn} emerging from the full set Umn∪VmnU_{mn}\cup V_{mn}, explicitly accounting for ties within UmnU_{mn} and VmnV_{mn} (davidson_luce, davidson_luce_map). Let πi>0\pi_i>0 be strengths and δt>0\delta_t>0 be tie-prevalence parameters with δ1≡1\delta_1\equiv 1. For a comparison set SS and tie order tt, define gt(T):=(∏i∈Tπi)1/tg_t(T):=\left(\prod_{i\in T}\pi_i\right)^{1/t} and

Z(S):=∑t′=1min⁡(D,∣S∣)δt′⋅∑T⊆S∣T∣=t′gt′(T),\begin{aligned} Z(S) := {}& \sum_{t'=1}^{\min(D,|S|)} \delta_{t'} \\ &\quad \cdot \sum_{\substack{T\subseteq S\\|T|=t'}} g_{t'}(T), \end{aligned} (50)

where DD is the maximum tie order considered. Then, for an event (U,V)(U,V) with S=U∪VS=U\cup V and t=∣U∣t=|U|,

Pr⁡(U≻V∣S)=δt gt(U)Z(S).\Pr(U\succ V\mid S) = \frac{\delta_t\,g_t(U)}{Z(S)}. (51)
Bradley - Terry - Luce (BTL) setwise-choice construction.

BTL converts each winner i∈Umni\in U_{mn} into a Luce choice event from {i}∪Vmn\{i\}\cup V_{mn}, with choice probability Pr⁡(i∣{i}∪V)=πi/(πi+∑j∈Vπj)\Pr(i\mid \{i\}\cup V)=\pi_i/(\pi_i+\sum_{j\in V}\pi_j) (bradley_terry_luce, bradley_terry_luce_map). Equivalently, for an event (U,V)(U,V) the BTL likelihood factorizes as

Pr⁡(U≻V)=∏i∈Uπiπi+∑j∈Vπj.\Pr(U\succ V) = \prod_{i\in U}\frac{\pi_i}{\pi_i+\sum_{j\in V}\pi_j}. (52)
Algorithm 4. MM algorithm for PL/BT on the pairwise win matrix
Require Pairwise win matrix W∈R+L×LW\in\mathbb{R}_+^{L\times L}, iterations TT
Ensure Strengths π^∈R+L\hat{\pi}\in\mathbb{R}_+^L (normalized)
wi←∑jWijw_i\gets \sum_j W_{ij} (total wins); nij←Wij+Wjin_{ij}\gets W_{ij}+W_{ji} (total comparisons)
Initialize πi∝wi\pi_i\propto w_i and normalize ∑iπi=1\sum_i\pi_i=1
For t=1t=1 to TT
For i=1i=1 to LL
di←∑j≠i: nij>0nijπi+πjd_i \gets \sum_{j\neq i:\,n_{ij}>0} \frac{n_{ij}}{\pi_i+\pi_j}
πi←wi/di\pi_i \gets w_i / d_i
EndFor
Normalize π\pi to sum to 11
EndFor
Return π\pi
Algorithm 5. Setwise event extraction and Luce-family estimation (Davidson - Luce / BTL)
Require R∈{0,1}L×M×NR\in\{0,1\}^{L\times M\times N}; model type ∈{Davidson–Luce,BTL}\in\{\text{Davidson--Luce},\text{BTL}\}; optional prior on log-strengths; max iterations TT
Ensure Strength scores π^∈R+L\hat{\pi}\in\mathbb{R}_+^L
Build events E←{(Umn,Vmn):0<∣Umn∣<L}\mathcal{E}\gets\{(U_{mn},V_{mn}) : 0<|U_{mn}|<L\}
Parameterize πi=exp⁡(θi)\pi_i=\exp(\theta_i) with centered θ\theta for identifiability
Define the event log-likelihood ∑(U,V)∈Elog⁡p(U≻V∣θ)\sum_{(U,V)\in\mathcal{E}}\log p(U\succ V\mid \theta) for the chosen model
Add prior penalty on θ\theta for MAP (or 00 for ML)
Optimize with L-BFGS for up to TT iterations and return π^\hat{\pi}

Item Response Theory (IRT) Methods

Scorio includes several IRT-inspired ranking methods that treat each model as an "examinee" with a latent ability and each question as an "item" with latent parameters (e.g., difficulty). We use IRT primarily as a ranking model: we estimate abilities {θl}l=1L\{\theta_l\}_{l=1}^L and rank models by θl\theta_l (larger is better), using rank_scores\texttt{rank\_scores} for tie-aware rank variants.

Data and binomial reduction.

Our raw observations are binary trial outcomes Rlmn∈{0,1}R_{lmn}\in\{0,1\} for model l∈{1,…,L}l\in\{1,\dots,L\}, question m∈{1,…,M}m\in\{1,\dots,M\}, and trial n∈{1,…,N}n\in\{1,\dots,N\}. When trials are i.i.d. conditional on parameters, the sufficient statistic for an item-model pair is the correct-count

klm:=∑n=1NRlmn∈{0,1,…,N},k_{lm} := \sum_{n=1}^N R_{lmn} \in \{0,1,\dots,N\}, (53)

so that likelihood-based IRT estimation can be written as a binomial-response model [20, 21].

Rasch (1PL).

The Rasch model [18] assumes a single item parameter (difficulty bmb_m):

klm∼Binomial ⁣(N,  σ(θl−bm)),k_{lm} \sim \mathrm{Binomial}\!\left(N,\;\sigma(\theta_l - b_m)\right), (54)

where σ(x)=1/(1+e−x)\sigma(x)=1/(1+e^{-x}). The model is invariant to global shifts (θ,b)↦(θ+c,b+c)(\theta,b)\mapsto(\theta+c,b+c), so we impose an identifiability constraint by centering item difficulties (e.g., ∑mbm=0\sum_m b_m = 0).

2PL and 3PL.

The 2PL model [19] adds an item discrimination parameter am>0a_m>0:

klm∼Binomial ⁣(N,  σ ⁣(am(θl−bm))).k_{lm} \sim \mathrm{Binomial}\!\left(N,\;\sigma\!\big(a_m(\theta_l - b_m)\big)\right). (55)

The 3PL model further adds a pseudo-guessing parameter cmc_m:

klm∼Binomial ⁣(N,  plm),plm:=cm+(1−cm)σ ⁣(am(θl−bm)).\begin{split} k_{lm} &\sim \mathrm{Binomial}\!\left(N,\;p_{lm}\right),\\ p_{lm} &:= c_m + (1-c_m)\sigma\!\big(a_m(\theta_l - b_m)\big). \end{split} (56)

In our implementation, we constrain ama_m via a log-parameterization and keep cmc_m in a bounded range (or optionally fix cmc_m to a known chance level).

Estimation variants used in Scorio.
  • JMLE / MLE (rasch, rasch_2pl, rasch_3pl): optimize the joint log-likelihood over θ\theta and item parameters.

  • MAP (rasch_map, rasch_2pl_map, rasch_3pl_map): add a prior penalty on abilities, typically Gaussian, as in Bayes modal estimation [40].

  • MML + EAP (rasch_mml): integrate out abilities under a population model (we use a standard normal prior), fit item parameters by EM, then compute EAP ability estimates [82, 41].

  • Credible/LB scoring (rasch_mml_credible): rank by a posterior quantile of θl\theta_l (e.g., a lower bound), which yields a conservative, uncertainty-aware ranking.

  • Dynamic IRT (dynamic_irt): a longitudinal extension that allows per-model trends across trials [22, 23].

Algorithm 6. Binomial xPL IRT (JMLE/MAP) for ranking
Require Response tensor R∈{0,1}L×M×NR\in\{0,1\}^{L\times M\times N}; model type ∈{1PL,2PL,3PL}\in\{\text{1PL},\text{2PL},\text{3PL}\}; optional ability prior p(θ)p(\theta); max iterations TT
Ensure Ability scores θ^∈RL\hat{\theta}\in\mathbb{R}^{L} and optional item parameters
Compute counts klm←∑n=1NRlmnk_{lm}\gets \sum_{n=1}^N R_{lmn} and set n←Nn\gets N
Initialize θ\theta from per-model accuracy; initialize bb from per-item solve rate; set am←1a_m\gets 1 (2PL/3PL); set cm←0.25c_m\gets 0.25 (3PL)
Define plm(θ,b,a,c)p_{lm}(\theta,b,a,c) according to the chosen xPL link
Define the binomial log-likelihood ℓ(k;n,p)←klog⁡p+(n−k)log⁡(1−p)\ell(k;n,p)\gets k\log p + (n-k)\log(1-p)
Define objective (negative log posterior)
\[
aligned
L(,b,a,c) =& -_l,m (k_lm; n, p_lm)
&- p().
aligned
\]
Set log⁡p(θ)=0\log p(\theta)=0 for pure MLE.
Impose identifiability at each iteration by centering item difficulties: b←b−1M∑mbmb \leftarrow b - \frac{1}{M}\sum_m b_m
Optimize L\mathcal{L} with a quasi-Newton method (e.g., L-BFGS) for up to TT iterations
Return θ^\hat{\theta} as scores (larger is better) and optionally b^\hat{b}, a^\hat{a}, c^\hat{c}
Algorithm 7. Rasch MML (EM + quadrature) with EAP and posterior-quantile scoring
Require Counts k∈{0,…,N}L×Mk\in\{0,\dots,N\}^{L\times M}; trials NN; quadrature points {θq,wq}q=1Q\{\theta_q,w_q\}_{q=1}^Q; EM iterations SS
Ensure EAP scores θ^EAP\hat{\theta}^{\mathrm{EAP}} (or quantile scores) and item difficulties b^\hat{b}
Initialize item difficulties bb from per-item solve rates and center bb
For s=1s=1 to SS
E-step: compute log⁡p(kl∣θq,b)\log p(k_l \mid \theta_q,b) for each model ll and quadrature point qq
Compute posterior weights wlq∝exp⁡(log⁡p(kl∣θq,b)) wqw_{lq}\propto \exp(\log p(k_l\mid \theta_q,b))\,w_q and normalize over qq
Define ℓ(k;n,p)←klog⁡p+(n−k)log⁡(1−p)\ell(k;n,p)\gets k\log p + (n-k)\log(1-p)
M-step: for each item mm, update bmb_m by minimizing
\[
-_l,q w_lq(k_lm;N,(_q-b_m)).
\]
Center bb
EndFor
Recompute posterior weights wlqw_{lq} under final bb
Compute EAP scores: θ^lEAP←∑qwlqθq\hat{\theta}^{\mathrm{EAP}}_l \gets \sum_q w_{lq}\theta_q
(Optional) Compute quantile score Qα(θl∣k)Q_\alpha(\theta_l\mid k) from the discrete posterior CDF (used by rasch_mml_credible)
Return scores and b^\hat{b}
Algorithm 8. Dynamic IRT growth model (logistic longitudinal Rasch)
Require Response tensor R∈{0,1}L×M×NR\in\{0,1\}^{L\times M\times N}; normalized time grid tn∈[0,1]t_n\in[0,1]; max iterations TT
Ensure Baseline abilities θ^0∈RL\hat{\theta}_0\in\mathbb{R}^L, slopes θ^1∈RL\hat{\theta}_1\in\mathbb{R}^L, and item difficulties b^∈RM\hat{b}\in\mathbb{R}^M
Fit the longitudinal model P(Rlmn=1)=σ(θ0,l+θ1,ltn−bm)P(R_{lmn}=1)=\sigma(\theta_{0,l}+\theta_{1,l}t_n-b_m) by maximizing the Bernoulli likelihood over all (l,m,n)(l,m,n)
Add weak regularization on slopes (e.g., ∥θ1∥22\|\theta_1\|_2^2) to avoid overfitting i.i.d. sampling noise
Center bb for identifiability
Optimize with a quasi-Newton method (e.g., L-BFGS) for up to TT iterations
Return θ^0\hat{\theta}_0 as ranking scores and optionally θ^1,b^\hat{\theta}_1,\hat{b}

Graph and Spectral Methods

These methods operate on the pairwise comparison graph derived from the win/tie counts (Wij,Tij)(W_{ij},T_{ij}) defined in section 2.2. A common derived quantity is the empirical tied-split win probability

P^i≻j:=Wij+12TijWij+Wji+Tij,P^i≻i:=12.\widehat{P}_{i\succ j} := \frac{W_{ij}+\tfrac{1}{2}T_{ij}}{W_{ij}+W_{ji}+T_{ij}},\qquad \widehat{P}_{i\succ i}:=\tfrac{1}{2}. (57)

In our fully observed benchmark setting, Wij+Wji+Tij=MNW_{ij}+W_{ji}+T_{ij}=MN for all i≠ji\neq j (section 2.2), so P^i≻j\widehat{P}_{i\succ j} is a simple rescaling of aggregated counts.

PageRank.

We build a directed weighted graph where an edge from jj to ii has weight P^i≻j\widehat{P}_{i\succ j} (interpreting "losers link to winners"), then form a column-stochastic transition matrix PP by normalizing each column:

Pij:=P^i≻j∑k≠jP^k≻j(i≠j),P_{ij} := \frac{\widehat{P}_{i\succ j}}{\sum_{k\neq j}\widehat{P}_{k\succ j}}\quad (i\neq j), (58)

with the standard dangling-node convention of a uniform column if the denominator is zero. PageRank scores r∈ΔL−1r\in\Delta^{L-1} solve

r=d P r+(1−d) 1L1,r = d\,P\,r + (1-d)\,\tfrac{1}{L}\mathbf{1}, (59)

where d∈(0,1)d\in(0,1) is the damping factor and 1\mathbf{1} is the all-ones vector [27]. This corresponds to pagerank in Scorio.

Spectral (eigenvector centrality).

We form the nonnegative matrix WW with off-diagonal entries Wij:=P^i≻jW_{ij}:=\widehat{P}_{i\succ j} and set the diagonal to the row sum Wii:=∑j≠iWijW_{ii}:=\sum_{j\neq i}W_{ij} (a self-loop that makes the matrix diagonally dominant). The spectral score vector is the principal right eigenvector v≥0v\ge 0 of WW, normalized to ∑ivi=1\sum_i v_i=1. This corresponds to spectral in Scorio.

Rank Centrality.

Rank Centrality [28] constructs a random walk on the comparison graph whose transition probabilities prefer moving from a model to those that beat it. Let dmaxd_{\mathrm{max}} be the maximum (undirected) degree of the comparison graph (in our benchmark setting dmax=L−1d_{\mathrm{max}}=L-1). Define a row-stochastic matrix

Pij:=1dmaxP^j≻i(i≠j),Pii:=1−∑j≠iPij.\begin{aligned} P_{ij} &:= \frac{1}{d_{\mathrm{max}}}\widehat{P}_{j\succ i}\quad (i\neq j),\\ P_{ii} &:= 1-\sum_{j\neq i}P_{ij}. \end{aligned} (60)

The stationary distribution π\pi of PP is used as the score vector (larger πi\pi_i is better). This corresponds to rank_centrality in Scorio.

α\alpha-Rank.

α\alpha-Rank [31] ranks strategies via evolutionary dynamics by constructing a Markov chain over models using fixation probabilities in a finite population. In our constant-sum binary evaluation setting, we treat P^i≻j\widehat{P}_{i\succ j} as the payoff to strategy ii against jj (so the per-match payoff sum is 11 when ties are split as 12\tfrac{1}{2}). For population size m≥2m\ge 2 and selection intensity α≥0\alpha\ge 0, the (constant-sum) fixation probability of a mutant rr in a resident population ss is

ρr,s:={1−exp⁡(−u)1−exp⁡(−mu)u≠0,1mu=0,\rho_{r,s} := \begin{cases} \frac{1-\exp(-u)}{1-\exp(-m u)} & u\neq 0,\\ \frac{1}{m} & u=0, \end{cases} (61)
u:=αmm−1(P^r≻s−12).u := \alpha\frac{m}{m-1}\left(\widehat{P}_{r\succ s}-\tfrac{1}{2}\right). (62)

The induced Markov chain on models has off-diagonal transitions Csr:=1L−1ρr,sC_{s r}:=\frac{1}{L-1}\rho_{r,s} and diagonal Css:=1−∑r≠sCsrC_{ss}:=1-\sum_{r\neq s}C_{sr}; the stationary distribution of CC is the α\alpha-Rank score vector. This corresponds to alpharank in Scorio.

Nash equilibrium mixture.

Following the use of Nash equilibria as evaluation summaries in symmetric zero-sum games [32], we define a zero-sum payoff matrix

Aij:=2P^i≻j−1,Aii:=0,A_{ij} := 2\widehat{P}_{i\succ j}-1,\qquad A_{ii}:=0, (63)

which is antisymmetric when P^\widehat{P}is derived from tied-split win rates. We compute a maximin mixed strategy x∈ΔL−1x\in\Delta^{L-1} (a Nash equilibrium strategy for the row player)

x∈arg⁡max⁡x∈ΔL−1min⁡y∈ΔL−1x⊤Ay,x \in \arg\max_{x\in\Delta^{L-1}} \min_{y\in\Delta^{L-1}} x^\top A y, (64)

via a standard linear program. To obtain a per-model evaluation score ("Nash averaging"), we then score each model by its expected performance against the equilibrium mixture opponent:

si:=∑j=1LP^i≻j xj∈[0,1],s_i := \sum_{j=1}^L \widehat{P}_{i\succ j}\,x_j \in [0,1], (65)

and rank models by ss (higher is better). We additionally report the equilibrium mixture xx as a strategic summary of the meta-game when needed. This corresponds to nash in Scorio.

Seriation-based Methods

SerialRank.

SerialRank [30] is a spectral seriation method that constructs a similarity graph from a skew-symmetric comparison matrix. From pairwise counts (W,T)(W,T), define

Cij:=Wij−WjiWij+Wji+Tij∈[−1,1],Cii:=0,\begin{aligned} C_{ij}&:=\frac{W_{ij}-W_{ji}}{W_{ij}+W_{ji}+T_{ij}}\in[-1,1],\\ C_{ii}&:=0, \end{aligned} (66)

so that Cij>0C_{ij}>0 indicates ii tends to beat jj (and CC is skew-symmetric). SerialRank forms the similarity matrix

S:=12(L 11⊤+CC⊤),S := \tfrac{1}{2}\left(L\,\mathbf{1}\mathbf{1}^\top + C C^\top\right), (67)

then computes the graph Laplacian LS:=diag(S1)−SL_S:=\mathrm{diag}(S\mathbf{1})-S. The ordering is given by sorting a Fiedler vector (the eigenvector associated with the second-smallest eigenvalue of LSL_S), with the sign chosen to best agree with the observed comparisons. This corresponds to serial_rank in Scorio.

Hodge-theoretic Methods

HodgeRank.

HodgeRank [29] interprets pairwise comparisons as a skew-symmetric edge flow on a graph and recovers global scores by least squares. Using the tied-split probabilities from appendix J.1.9, define the observed edge flow

Y‾ij:=P^j≻i−P^i≻j=Wji−WijWij+Wji+Tij,Y‾ii:=0,\begin{aligned} \overline{Y}_{ij} &:= \widehat{P}_{j\succ i}-\widehat{P}_{i\succ j} = \frac{W_{ji}-W_{ij}}{W_{ij}+W_{ji}+T_{ij}},\\ \overline{Y}_{ii} &:=0, \end{aligned} (68)

and choose symmetric edge weights wijw_{ij} (e.g., the total number of comparisons on edge (i,j)(i,j)). HodgeRank solves the weighted least-squares problem

s⋆∈arg⁡min⁡s∈RL  ∑i<jwij((sj−si)−Y‾ij)2=arg⁡min⁡s∥grad(s)−Y‾∥2,w2,\begin{aligned} s^\star &\in \arg\min_{s\in\mathbb{R}^L}\;\sum_{i<j} w_{ij}\left((s_j-s_i)-\overline{Y}_{ij}\right)^2\\ &= \arg\min_s \|\mathrm{grad}(s)-\overline{Y}\|_{2,w}^2, \end{aligned} (69)

which reduces to a weighted graph Laplacian system; we compute the minimum-norm solution via the Moore - Penrose pseudoinverse and rank by s⋆s^\star (higher is better). This corresponds to hodge_rank in Scorio.

Ranking Method APIs and Hyperparameters

We evaluate the ranking methods described in appendix J.1. Each method maps the trial outcome tensor R∈{0,1}L×M×NR\in\{0,1\}^{L\times M\times N} (and, where applicable, an optional prior tensor R0R_0) to a ranking over the LL models. For reproducibility, we list the exact API identifiers and argument values used in our experiments; None denotes an unset optional argument.

Metrics.
  • avg

  • pass_at_k_2 (k=2)

  • pass_hat_k_2 (k=2)

  • mg_pass_at_k_2 (k=2)

  • bayes (R0=None, quantile=None)

  • bayes_greedy (R0=R0, quantile=None)

  • bayes_ci (R0=None, quantile=0.05)

  • inverse_difficulty (return_scores=false, clip_range=[0.01, 0.99])

Pairwise rating.
  • elo_tie_skip (K=0.05, initial_rating=1500.0, tie_handling=skip)

  • elo_tie_draw (K=0.05, initial_rating=1500.0, tie_handling=draw)

  • elo_tie_correct_draw_only (K=0.05, initial_rating=1500.0, tie_handling=correct_draw_only)

  • glicko_tie_skip (initial_rating=1500.0, initial_rd=350.0, c=0.0, rd_max=350.0, tie_handling=skip, return_deviation=false)

  • glicko_tie_draw (initial_rating=1500.0, initial_rd=350.0, c=0.0, rd_max=350.0, tie_handling=draw, return_deviation=false)

  • glicko_tie_correct_draw_only (initial_rating=1500.0, initial_rd=350.0, c=0.0, rd_max=350.0, tie_handling=correct_draw_only, return_deviation=false)

  • trueskill (mu_initial=25.0, sigma_initial=8.333333333333334, beta=4.166666666666667, tau=0.00333333333)

Probabilistic comparisons.
  • bradley_terry (return_scores=false, max_iter=500)

  • bradley_terry_map (prior=1.0, max_iter=500)

  • bradley_terry_davidson (return_scores=false, max_iter=500)

  • bradley_terry_davidson_map (prior=1.0, max_iter=500)

  • rao_kupper (tie_strength=1.1, max_iter=500)

  • rao_kupper_map (tie_strength=1.1, prior=1.0, max_iter=500)

  • thompson (n_samples=10000, prior_alpha=1.0, prior_beta=1.0, seed=42)

  • bayesian_mcmc (n_samples=5000, burnin=1000, prior_var=1.0, seed=42)

  • plackett_luce (return_scores=false, max_iter=500, tol=1e-08)

  • plackett_luce_map (prior=1.0, max_iter=500)

  • bradley_terry_luce (return_scores=false, max_iter=500)

  • bradley_terry_luce_map (prior=1.0, max_iter=500)

Voting rules.
  • borda (return_scores=false)

  • copeland (return_scores=false)

  • win_rate (return_scores=false)

  • minimax_variant_margin_tie_ignore (variant=margin, tie_policy=ignore)

  • minimax_variant_margin_tie_half (variant=margin, tie_policy=half)

  • minimax_ variant_ winning_ votes_ tie_ ignore (variant=winning_ votes, tie_ policy=ignore)

  • minimax_ variant_ winning_ votes_ tie_ half (variant=winning_ votes, tie_ policy=half)

  • schulze_tie_ignore (tie_policy=ignore)

  • schulze_tie_half (tie_policy=half)

  • ranked_ pairs_ strength_ margin_ tie_ ignore (strength=margin, tie_ policy=ignore)

  • ranked_ pairs_ strength_ margin_ tie_ half (strength=margin, tie_ policy=half)

  • ranked_ pairs_ strength_ winning_ votes_ tie_ ignore (strength=winning_ votes, tie_ policy=ignore)

  • ranked_ pairs_ strength_ winning_ votes_ tie_ half (strength=winning_ votes, tie_ policy=half)

  • kemeny_young_tie_ignore (tie_policy=ignore, time_limit=None)

  • kemeny_young_tie_half (tie_policy=half, time_limit=None)

  • nanson_rank_ties_average (rank_ties=average)

  • nanson_rank_ties_max (rank_ties=max)

  • baldwin_rank_ties_average (rank_ties=average)

  • baldwin_rank_ties_max (rank_ties=max)

  • majority_judgment (return_scores=false)

IRT.
  • rasch (return_scores=false, max_iter=500, return_item_params=false)

  • rasch_map (prior=1.0, max_iter=500, return_item_params=false)

  • rasch_2pl (return_scores=false, max_iter=500, return_item_params=false)

  • rasch_2pl_map (prior=1.0, max_iter=500, return_item_params=false)

  • rasch_3pl (return_scores=false, max_iter=500, fix_guessing=None, return_item_params=false)

  • rasch_3pl_map (prior=1.0, max_iter=500, fix_guessing=None, return_item_params=false)

  • rasch_mml (return_scores=false, max_iter=100, em_iter=20, n_quadrature=21, return_item_params=false)

  • rasch_mml_credible (quantile=0.05, max_iter=100, em_iter=20, n_quadrature=21)

  • dynamic_irt_linear (variant=linear, max_iter=500, return_item_params=false)

  • dynamic_irt_growth (variant=growth, max_iter=500, return_item_params=false)

Graph/game.
  • pagerank (damping=0.85, max_iter=100, tol=1e-12)

  • spectral (max_iter=10000, tol=1e-12)

  • alpharank (alpha=1.0, population_size=50, max_iter=100000, tol=1e-12)

  • nash_vs_equilibrium (n_iter=100, temperature=0.1, solver=lp, score_type=vs_equilibrium, return_equilibrium=false)

  • nash_ advantage_ vs_ equilibrium (n_ iter=100, temperature=0.1, solver=lp, score_ type=advantage_ vs_ equilibrium, return_ equilibrium=false)

  • rank_centrality_tie_ignore (tie_handling=ignore, smoothing=0.0, teleport=0.0, max_iter=10000, tol=1e-12)

  • rank_centrality_tie_half (tie_handling=half, smoothing=0.0, teleport=0.0, max_iter=10000, tol=1e-12)

  • serial_rank_prob_diff (comparison=prob_diff)

  • serial_rank_sign (comparison=sign)

  • hodge_rank_binary_total (pairwise_stat=binary, weight_method=total, return_diagnostics=false)

  • hodge_rank_binary_decisive (pairwise_stat=binary, weight_method=decisive, return_diagnostics=false)

  • hodge_rank_binary_uniform (pairwise_stat=binary, weight_method=uniform, return_diagnostics=false)

  • hodge_rank_log_odds_total (pairwise_stat=log_odds, weight_method=total, epsilon=0.5, return_diagnostics=false)

  • hodge_rank_log_odds_decisive (pairwise_stat=log_odds, weight_method=decisive, epsilon=0.5, return_diagnostics=false)

  • hodge_rank_log_odds_uniform (pairwise_stat=log_odds, weight_method=uniform, epsilon=0.5, return_diagnostics=false)

References

  1. Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, Wojciech Zaremba. Evaluating Large Language Models Trained on Code. misc. 2021. https://arxiv.org/abs/2107.03374
  2. Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou. Self-Consistency Improves Chain of Thought Reasoning in Language Models. International Conference on Learning Representations. 2023. https://openreview.net/forum?id=1PL1NIMMrw
  3. Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. misc. 2024. https://arxiv.org/abs/2408.03314
  4. Zeng, Zhiyuan, Chen, Qingyuan, Yin, Zhangyue, Zhou, Yunhua, Qiu, Xipeng. Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. https://aclanthology.org/2025.acl-long.232/
  5. Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, Dario Amodei. Deep Reinforcement Learning from Human Preferences. Advances in Neural Information Processing Systems. 2017. https://papers.nips.cc/paper/7017-deep-reinforcement-learning-from-human-preferences
  6. Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, Chelsea Finn. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. Advances in Neural Information Processing Systems. 2023. https://papers.nips.cc/paper_files/paper/2023/hash/a85b405ed65c6477a4fe8302b5e06ce7-Abstract-Conference.html
  7. Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E. Gonzalez, Ion Stoica. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. Proceedings of the 41st International Conference on Machine Learning. 2024. https://proceedings.mlr.press/v235/chiang24b.html
  8. Siavash Ameli, Siyuan Zhuang, Ion Stoica, Michael W. Mahoney. A Statistical Framework for Ranking LLM-based Chatbots. International Conference on Learning Representations. 2025. https://openreview.net/forum?id=rAoEub6Nw2
  9. Hongli Zhou, Hui Huang, Ziqing Zhao, Lvyuan Han, Huicheng Wang, Kehai Chen, Muyun Yang, Wei Bao, Jian Dong, Bing Xu, Conghui Zhu, Hailong Cao, Tiejun Zhao. Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory. misc. 2025. https://arxiv.org/abs/2505.15055
  10. Bradley, Ralph Allan, Terry, Milton E.. Rank Analysis of Incomplete Block Designs: The Method of Paired Comparisons. Biometrika. 1952. https://doi.org/10.1093/biomet/39.3-4.324
  11. Elo, Arpad E.. The Rating of Chessplayers, Past and Present. Arco Publishing. 1978. https://archive.org/details/ratingofchesspla0000eloa
  12. Glickman, Mark E.. Parameter Estimation in Large Dynamic Paired Comparison Experiments. Journal of the Royal Statistical Society: Series C (Applied Statistics). 1999. https://doi.org/10.1111/1467-9876.00159
  13. de Borda, Jean-Charles. M\'emoire sur les \'elections au scrutin. misc. 1781. https://webusers.imj-prg.fr/ alexandre.guilbaud/LX2U1/Borda_Memoire_sur_les_elections_au_scrutin_MARS_1781_extrait.pdf
  14. Condorcet, Marie Jean Antoine Nicolas Caritat, Marquis de. Essai sur l'application de l'analyse `a la probabilit\'e des d\'ecisions rendues `a la pluralit\'e des voix. Imprimerie Royale. 1785. https://archive.org/details/bub_gb_RzAVAAAAQAAJ
  15. Arrow, Kenneth J.. Social Choice and Individual Values. John Wiley & Sons. 1951.
  16. Hariri, Mohsen, Samandar, Amirhossein, Hinczewski, Michael, Chaudhary, Vipin. Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation. Proceedings of the 14th International Conference on Learning Representations (ICLR 2026). 2026. https://openreview.net/forum?id=PTXi3Ef4sT
  17. Gotou, Takumi, Nagata, Ryo, Mita, Masato, Hanawa, Kazuaki. Taking the Correction Difficulty into Account in Grammatical Error Correction Evaluation. Proceedings of the 28th International Conference on Computational Linguistics. 2020. https://aclanthology.org/2020.coling-main.188/
  18. Rasch, Georg. Probabilistic Models for Some Intelligence and Attainment Tests. Danish Institute for Educational Research. 1960. https://archive.org/details/probabilisticmod0000rasc
  19. Birnbaum, Allan. Some Latent Trait Models and Their Use in Inferring an Examinee's Ability. Statistical Theories of Mental Test Scores. 1968. https://faculty.ucmerced.edu/jvevea/classes/290_21/readings/week%209/Birnbaum.pdf
  20. McCullagh, P., Nelder, J. A.. Generalized Linear Models. Springer. 1989. https://doi.org/10.1007/978-1-4899-3242-6
  21. Explanatory Item Response Models. Springer. 2004. https://doi.org/10.1007/978-1-4757-3990-9
  22. Verhelst, Norman D., Glas, Cees A. W.. A Dynamic Generalization of the Rasch Model. Psychometrika. 1993. https://doi.org/10.1007/BF02294648
  23. Wang, Chun, Nydick, Steven W.. On Longitudinal Item Response Theory Models: A Didactic. Journal of Educational and Behavioral Statistics. 2020. https://doi.org/10.3102/1076998619882026
  24. Rao, P. V., Kupper, L. L.. Ties in Paired-Comparison Experiments: A Generalization of the Bradley - Terry Model. Journal of the American Statistical Association. 1967. https://doi.org/10.1080/01621459.1967.10482901
  25. Davidson, Roger R.. On Extending the Bradley - Terry Model to Accommodate Ties in Paired Comparison Experiments. Journal of the American Statistical Association. 1970. https://doi.org/10.1080/01621459.1970.10481082
  26. Handbook of Computational Social Choice. Cambridge University Press. 2016. https://doi.org/10.1017/CBO9781107446984
  27. Lawrence Page, Sergey Brin, Rajeev Motwani, Terry Winograd. The PageRank Citation Ranking: Bringing Order to the Web. techreport. 1999. https://ilpubs.stanford.edu/422/1/1999-66.pdf
  28. Negahban, Sahand, Oh, Sewoong, Shah, Devavrat. Rank Centrality: Ranking from Pairwise Comparisons. Operations Research. 2017. https://doi.org/10.1287/opre.2016.1534
  29. Xiaoye Jiang, Lek-Heng Lim, Yuan Yao, Yinyu Ye. Statistical Ranking and Combinatorial Hodge Theory. Mathematical Programming. 2011. https://doi.org/10.1007/s10107-010-0419-x
  30. Fogel, Fajwel, d'Aspremont, Alexandre, Vojnovic, Milan. Spectral Ranking using Seriation. Journal of Machine Learning Research. 2016. https://jmlr.org/papers/v17/16-035.html
  31. Omidshafiei, Shayegan, Papadimitriou, Christos, Piliouras, Georgios, Tuyls, Karl, Rowland, Mark, Lespiau, Jean-Baptiste, Czarnecki, Wojciech M., P\'erolat, Julien, Munos, R\'emi. -Rank: Multi-Agent Evaluation by Evolution. Scientific Reports. 2019. https://doi.org/10.1038/s41598-019-45619-9
  32. Balduzzi, David, Garnelo, Marta, Bachrach, Yoram, Czarnecki, Wojciech, P\'erolat, Julien, Jaderberg, Max, Graepel, Thore. Open-ended Learning in Symmetric Zero-sum Games. Proceedings of the 36th International Conference on Machine Learning. 2019. https://proceedings.mlr.press/v97/balduzzi19a.html
  33. Ralf Herbrich, Tom Minka, Thore Graepel. TrueSkill: A Bayesian Skill Rating System. Advances in Neural Information Processing Systems. 2006. https://papers.neurips.cc/paper/3079-trueskilltm-a-bayesian-skill-rating-system.pdf
  34. Plackett, R. L.. The Analysis of Permutations. Applied Statistics. 1975. https://doi.org/10.2307/2346567
  35. Luce, R. Duncan. Individual Choice Behavior: A Theoretical Analysis. John Wiley & Sons. 1959. https://archive.org/details/individualchoice0000luce
  36. David Firth, Ioannis Kosmidis, Heather Turner. Davidson - Luce Model for Multi-item Choice with Ties. misc. 2019. https://arxiv.org/abs/1909.07123
  37. Hunter, David R.. MM Algorithms for Generalized Bradley - Terry Models. The Annals of Statistics. 2004. https://doi.org/10.1214/aos/1079120141
  38. Gelman, Andrew, Carlin, John B., Stern, Hal S., Dunson, David B., Vehtari, Aki, Rubin, Donald B.. Bayesian Data Analysis. CRC Press. 2013. https://doi.org/10.1201/b16018
  39. Caron, Fran, Doucet, Arnaud. Efficient Bayesian Inference for Generalized Bradley - Terry Models. Journal of Computational and Graphical Statistics. 2012. https://doi.org/10.1080/10618600.2012.638220
  40. Mislevy, Robert J.. Bayes Modal Estimation in Item Response Models. Psychometrika. 1986. https://doi.org/10.1007/BF02293979
  41. Chen, Ssu-Kuang, Hou, Liling, Dodd, Barbara G.. A Comparison of Maximum Likelihood Estimation and Expected a Posteriori Estimation in CAT Using the Partial Credit Model. Educational and Psychological Measurement. 1998. https://doi.org/10.1177/0013164498058004002
  42. Junnan Liu, Hongwei Liu, Linchen Xiao, Ziyi Wang, Kuikun Liu, Songyang Gao, Wenwei Zhang, Songyang Zhang, Kai Chen. Are Your LLMs Capable of Stable Reasoning?. Findings of the Association for Computational Linguistics: ACL 2025. 2025. https://aclanthology.org/2025.findings-acl.905/
  43. Rodriguez, Pedro, Barrow, Joe, Hoyle, Alexander, Lalor, John P., Jia, Robin, Boyd-Graber, Jordan. Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. https://aclanthology.org/2021.acl-long.346/
  44. Wolf, Thomas, Debut, Lysandre, Sanh, Victor, Chaumond, Julien, Delangue, Clement, Moi, Anthony, Cistac, Pierric, Rault, Tim, Louf, R\'emi, Funtowicz, Morgan, others. Huggingface's transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771. 2019. https://arxiv.org/abs/1910.03771
  45. Sylvain Gugger, Lysandre Debut, Thomas Wolf, Philipp Schmid, Zachary Mueller, Sourab Mangrulkar, Marc Sun, Benjamin Bossan. Accelerate: Training and inference at scale made simple, efficient and adaptable.. misc. 2022. https://github.com/huggingface/accelerate
  46. Dao, Tri. Flashattention-2: Faster attention with better parallelism and work partitioning. arXiv preprint arXiv:2307.08691. 2023. https://arxiv.org/abs/2307.08691
  47. Zhang, Tianyi, Hariri, Mohsen, Zhong, Shaochen, Chaudhary, Vipin, Sui, Yang, Hu, Xia, Shrivastava, Anshumali. 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11). Advances in Neural Information Processing Systems. 2025. https://arxiv.org/abs/2504.11651
  48. Mathematical Association of America. American Invitational Mathematics Examination (AIME). misc. 2024. https://maa.org/maa-invitational-competitions/
  49. Mathematical Association of America. American Invitational Mathematics Examination (AIME). misc. 2025. https://maa.org/maa-invitational-competitions/
  50. Brown University Math Olympiad Organizers. Brown University Math Olympiad (BrUMO). misc. 2025. https://www.brumo.org/tournament-info
  51. Harvard - MIT Mathematics Tournament. HMMT February 2025 Archive (Problems and Solutions). misc. 2025. https://www.hmmt.org/www/archive/282
  52. NovaSky Team. Think Less, Achieve More: Cut Reasoning Costs by 50% Without Sacrificing Accuracy. misc. 2025. https://novasky-ai.github.io/posts/reduce-overthinking/
  53. Qwen Team. Qwen3 Technical Report. misc. 2025. https://arxiv.org/abs/2505.09388
  54. Guo, D., Yang, D., Zhang, H., Song, J., Wang, P., Zhu, Q., Xu, R., Zhang, R., Ma, S., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., Xue, B., Wang, B., Wu, B., Feng, B., Lu, C., Zhao, C., Deng, C., Ruan, C., Dai, D., Chen, D., Ji, D., Li, E., Lin, F., Dai, F., Luo, F., Hao, G., Chen, G., Li, G., Zhang, H., Xu, H., Ding, H., Gao, H., Qu, H., Li, H., Guo, J., Li, J., Chen, J., Yuan, J., Tu, J., Qiu, J., Li, J., Cai, J. L., Ni, J., Liang, J., Chen, J., Dong, K., Hu, K., You, K., Gao, K., Guan, K., Huang, K., Yu, K., Wang, L., Zhang, L., Zhao, L., Wang, L., Zhang, L., Xu, L., Xia, L., Zhang, M., Zhang, M., Tang, M., Zhou, M., Li, M., Wang, M., Li, M., Tian, N., Huang, P., Zhang, P., Wang, Q., Chen, Q., Du, Q., Ge, R., Zhang, R., Pan, R., Wang, R., Chen, R. J., Jin, R. L., Chen, R., Lu, S., Zhou, S., Chen, S., Ye, S., Wang, S., Yu, S., Zhou, S., Pan, S., Li, S. S., Zhou, S., Wu, S., Yun, T., Pei, T., Sun, T., Wang, T., Zeng, W., Liu, W., Liang, W., Gao, W., Yu, W., Zhang, W., Xiao, W. L., An, W., Liu, X., Wang, X., Chen, X., Nie, X., Cheng, X., Liu, X., Xie, X., Liu, X., Yang, X., Li, X., Su, X., Lin, X., Li, X. Q., Jin, X., Shen, X., Chen, X., Sun, X., Wang, X., Song, X., Zhou, X., Wang, X., Shan, X., Li, Y. K., Wang, Y. Q., Wei, Y. X., Zhang, Y., Xu, Y., Li, Y., Zhao, Y., Sun, Y., Wang, Y., Yu, Y., Zhang, Y., Shi, Y., Xiong, Y., He, Y., Piao, Y., Wang, Y., Tan, Y., Ma, Y., Liu, Y., Guo, Y., Ou, Y., Wang, Y., Gong, Y., Zou, Y., He, Y., Xiong, Y., Luo, Y., You, Y., Liu, Y., Zhou, Y., Zhu, Y. X., Huang, Y., Li, Y., Zheng, Y., Zhu, Y., Ma, Y., Tang, Y., Zha, Y., Yan, Y., Ren, Z. Z., Ren, Z., Sha, Z., Fu, Z., Xu, Z., Xie, Z., Zhang, Z., Hao, Z., Ma, Z., Yan, Z., Wu, Z., Gu, Z., Zhu, Z., Liu, Z., Li, Z., Xie, Z., Song, Z., Pan, Z., Huang, Z., Xu, Z., Zhang, Z., Zhang, Z.. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature. 2025. https://doi.org/10.1038/s41586-025-09422-z
  55. OpenAI. gpt-oss-120b & gpt-oss-20b Model Card. misc. 2025. https://arxiv.org/abs/2508.10925
  56. Yixin Ye, Zhen Huang, Yang Xiao, Ethan Chern, Shijie Xia, Pengfei Liu. LIMO: Less is More for Reasoning. Second Conference on Language Modeling. 2025. https://openreview.net/forum?id=T2TZ0RY4Zk
  57. LG AI Research. EXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modes. misc. 2025. https://arxiv.org/abs/2507.11407
  58. NVIDIA. OpenReasoning-Nemotron-1.5B. misc. 2025. https://huggingface.co/nvidia/OpenReasoning-Nemotron-1.5B
  59. Etash Guha, Ryan Marten, Sedrick Keh, Negin Raoof, Georgios Smyrnis, Hritik Bansal, Marianna Nezhurina, Jean Mercat, Trung Vu, Zayne Sprague, Ashima Suvarna, Benjamin Feuer, Liangyu Chen, Zaid Khan, Eric Frankel, Sachin Grover, Caroline Choi, Niklas Muennighoff, Shiye Su, Wanjia Zhao, John Yang, Shreyas Pimpalgaonkar, Kartik Sharma, Charlie Cheng-Jie Ji, Yichuan Deng, Sarah Pratt, Vivek Ramanujan, Jon Saad-Falcon, Jeffrey Li, Achal Dave, Alon Albalak, Kushal Arora, Blake Wulfe, Chinmay Hegde, Greg Durrett, Sewoong Oh, Mohit Bansal, Saadia Gabriel, Aditya Grover, Kai-Wei Chang, Vaishaal Shankar, Aaron Gokaslan, Mike A. Merrill, Tatsunori Hashimoto, Yejin Choi, Jenia Jitsev, Reinhard Heckel, Maheswaran Sathiamoorthy, Alexandros G. Dimakis, Ludwig Schmidt. OpenThoughts: Data Recipes for Reasoning Models. misc. 2025. https://arxiv.org/abs/2506.04178
  60. Marah Abdin, Sahaj Agarwal, Ahmed Awadallah, Vidhisha Balachandran, Harkirat Behl, Lingjiao Chen, Gustavo de Rosa, Suriya Gunasekar, Mojan Javaheripi, Neel Joshi, Piero Kauffmann, Yash Lara, Caio César Teodoro Mendes, Arindam Mitra, Besmira Nushi, Dimitris Papailiopoulos, Olli Saarikivi, Shital Shah, Vaishnavi Shrivastava, Vibhav Vineet, Yue Wu, Safoora Yousefi, Guoqing Zheng. Phi-4-reasoning Technical Report. misc. 2025. https://arxiv.org/abs/2504.21318
  61. Hugging Face. Open-R1: A Fully Open Reproduction of DeepSeek-R1. misc. 2025. https://github.com/huggingface/open-r1
  62. FuseAI. FuseO1-DeepSeekR1-QwQ-SkyT1-Flash-32B-Preview. misc. 2025. https://huggingface.co/FuseAI/FuseO1-DeepSeekR1-QwQ-SkyT1-Flash-32B-Preview
  63. Liang Wen, Yunke Cai, Fenrui Xiao, Xin He, Qi An, Zhenyu Duan, Yimin Du, Junchen Liu, Lifu Tang, Xiaowei Lv, Haosheng Zou, Yongchao Deng, Shousheng Jia, Xiangzheng Zhang. Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track). 2025. https://aclanthology.org/2025.acl-industry.24/
  64. Zihan Liu, Zhuolin Yang, Yang Chen, Chankyu Lee, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping. AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy. International Conference on Learning Representations. 2026. https://openreview.net/forum?id=IaEqjWXd1d
  65. NVIDIA. NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model. misc. 2025. https://arxiv.org/abs/2508.14444
  66. Bespoke Labs. Bespoke-Stratos: The Unreasonable Effectiveness of Reasoning Distillation. misc. 2025. https://www.bespokelabs.ai/blog/bespoke-stratos-the-unreasonable-effectiveness-of-reasoning-distillation
  67. Kwon, Woosuk, Li, Zhuohan, Zhuang, Siyuan, Sheng, Ying, Zheng, Lianmin, Yu, Cody Hao, Gonzalez, Joseph, Zhang, Hao, Stoica, Ion. Efficient Memory Management for Large Language Model Serving with PagedAttention. Proceedings of the 29th Symposium on Operating Systems Principles. 2023. https://arxiv.org/abs/2309.06180
  68. Kendall, M. G.. A New Measure of Rank Correlation. Biometrika. 1938. https://doi.org/10.1093/biomet/30.1-2.81
  69. Shunyu Yao, Noah Shinn, Pedram Razavi, Karthik Narasimhan. -bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. International Conference on Learning Representations. 2025. https://openreview.net/forum?id=roNSXZpUDN
  70. Thompson, William R.. On the Likelihood that One Unknown Probability Exceeds Another in View of the Evidence of Two Samples. Biometrika. 1933. https://doi.org/10.1093/biomet/25.3-4.285
  71. Russo, Daniel J., Van Roy, Benjamin, Kazerouni, Abbas, Osband, Ian, Wen, Zheng. A Tutorial on Thompson Sampling. Foundations and Trends in Machine Learning. 2018. https://doi.org/10.1561/2200000070
  72. Metropolis, Nicholas, Rosenbluth, Arianna W., Rosenbluth, Marshall N., Teller, Augusta H., Teller, Edward. Equation of State Calculations by Fast Computing Machines. The Journal of Chemical Physics. 1953. https://doi.org/10.1063/1.1699114
  73. Hastings, W. K.. Monte Carlo Sampling Methods Using Markov Chains and Their Applications. Biometrika. 1970. https://doi.org/10.1093/biomet/57.1.97
  74. Copeland, Arthur H.. A Reasonable Social Welfare Function. misc. 1951. https://bibbase.org/network/publication/copeland-areasonablesocialwelfarefunction-1951
  75. Schulze, Markus. A new monotonic, clone-independent, reversal symmetric, and Condorcet-consistent single-winner election method. Social Choice and Welfare. 2011. https://doi.org/10.1007/s00355-010-0475-4
  76. Tideman, T. N.. Independence of clones as a criterion for voting rules. Social Choice and Welfare. 1987. https://doi.org/10.1007/BF00433944
  77. Kemeny, John G.. Mathematics without Numbers. Daedalus. 1959. https://www.jstor.org/stable/20026581
  78. Young, H. P.. Extending Condorcet's rule. Journal of Economic Theory. 1977. https://doi.org/10.1016/0022-0531(77)90012-6
  79. Nanson, E. J.. Methods of Election. Transactions and Proceedings of the Royal Society of Victoria. 1883. https://www.biodiversitylibrary.org/itemdetails/106382
  80. Baldwin, J. M.. The technique of the Nanson preferential majority system of election. Proceedings of the Royal Society of Victoria, New Series. 1926. https://www.biodiversitylibrary.org/itemdetails/236582
  81. Balinski, Michel, Laraki, Rida. Majority Judgment: Measuring, Ranking, and Electing. The MIT Press. 2011. https://doi.org/10.7551/mitpress/9780262015134.001.0001
  82. Bock, R. Darrell, Aitkin, Murray. Marginal Maximum Likelihood Estimation of Item Parameters: Application of an EM Algorithm. Psychometrika. 1981. https://doi.org/10.1007/BF02293801