Post
Scorio Datasets for Test-Time Scaling
Scorio Trace, Lite, Math, and GPQA contain 1.4 million repeated reasoning attempts for evaluation, ranking, answer selection, and token-level confidence analysis.
2 items tagged with "Datasets"
Scorio Trace, Lite, Math, and GPQA contain 1.4 million repeated reasoning attempts for evaluation, ranking, answer selection, and token-level confidence analysis.
A principled Bayesian framework that replaces Pass@k with posterior estimates, credible intervals, and stable rankings for LLM evaluation