Rank models by inverse-difficulty-weighted per-question accuracy.
Each question is weighted by the reciprocal of its clipped global solve rate, upweighting hard questions; a model's score is the weighted average of its per-question accuracies.
Rank models by inverse-difficulty-weighted per-question accuracy.
Each question is weighted by the reciprocal of its clipped global solve rate, upweighting hard questions; a model's score is the weighted average of its per-question accuracies.