Scorio - v0.2.3
    Preparing search index...

    Evaluation starts with repeated outcomes for one model. The functions average over questions, so each question contributes equally. The data guide describes binary and categorical inputs. The API reference lists the exported metrics and their signatures.

    bayes(R, w?, R0?) returns [mean, std]. It adds one prior count per category, then adds the observed counts from R and any prior outcomes from R0. std is the posterior standard deviation of the mean performance across questions.

    avg(R, w?) also returns two numbers. Its first value is the observed weighted average; its second is a Bayesian uncertainty estimate rescaled to the average scale. It does not accept R0.

    import { avg, bayes } from "scorio/eval";

    const R = [[0, 1, 1, 0, 1], [1, 1, 0, 1, 1]];
    const [observed, observedStd] = avg(R);
    const [posterior, posteriorStd] = bayes(R);

    console.log(observed); // => 0.7
    console.log(posterior); // about 0.642857
    console.log(observedStd, posteriorStd);

    The difference between the two means is expected. Bayes@N includes the prior, which pulls estimates away from the extremes when few trials are available.

    The scalar metrics below answer different questions about a bank of completions. For the binary pass and majority families, k is an integer from 1 to the observed trial count N.

    Function Quantity
    passAtK(R, k) Probability at least one of k selected trials succeeds
    passHatK(R, k), unanimousAtK(R, k), gPassAtK(R, k) Probability all k selected trials succeed
    gPassAtKTau(R, k, tau) Probability of at least max(1, ceil(tau × k)) successes
    mgPassAtK(R, k) Average of generalized pass thresholds over the upper half of the threshold range
    majAtK(R, k) Probability a strict majority of the selected trials succeeds
    aucAtK(R, k) Normalized trapezoidal area under Pass@1 through Pass@k
    maxAtK(R, k, w?) Expected maximum rubric score in k selected trials
    import { passAtK, passHatK, majAtK } from "scorio/eval";

    const R = [[0, 1, 1, 0, 1], [1, 1, 0, 1, 1]];
    console.log(passAtK(R, 2)); // => 0.95
    console.log(passHatK(R, 2)); // => 0.45
    console.log(majAtK(R, 3)); // => 0.85

    majAtK counts correct outcomes. Actual voting over answer strings is handled by majorityVote in aggregation.

    The API also includes geomAtK, geomDsAtK, geoSpectrumAtK, geoSpectrumStarAtK, and thresholdSpectrumAtK for geometric and threshold blends. Their parameters and definitions are in the API reference.

    The *Ci functions return [mean, std, lo, hi]. The usual confidence default is 0.95, and the interval uses a normal approximation. Pass-family intervals clip to [0, 1] by default; bayesCi and avgCi accept optional bounds and do not clip unless you pass them.

    import { bayesCi, passAtK, passAtKCi } from "scorio/eval";

    const R = [[0, 1, 1, 0, 1], [1, 1, 0, 1, 1]];
    const [mean, std, lo, hi] = bayesCi(R, undefined, undefined, 0.95, [0, 1]);
    console.log(mean, std, lo, hi);

    console.log(passAtK(R, 2)); // empirical estimate: 0.95
    console.log(passAtKCi(R, 2)[0]); // posterior mean: about 0.839286

    The interval companion's mean can differ from the scalar point estimate. For passAtK, the point estimator samples the observed bank without replacement. passAtKCi instead summarizes the i.i.d. success quantity under a Beta posterior, with success and failure prior concentrations both 1 by default. Report which estimate you use.

    TailPass has its own profile and interval API. Read the TailPass guide when you need several success thresholds or want to choose a utility explicitly.