AI Model Index

Scoring System and Methodology

Inspect how raw benchmark rows become normalized, coverage-aware, confidence-weighted, reproducible AI model index scores.

Frequently asked questions

How does a raw benchmark row become an index score?

Rows are collected with their original identity, date, unit, and source context, reconciled to the exact model measured, normalized within declared boundaries, then combined with coverage-aware, confidence-weighted component weights. Formula versions and contribution details stay published, so any score can be recomputed.

Why not average every benchmark number into one score?

Percentages, Elo ratings, ranks, pass rates, cost, and speed measure different things on incompatible scales. The methodology keeps model-only, agent, price, and display-only evidence separate, requires multiple evaluator families for broad rankings, and limits how much one evaluator can dominate a composite.

Are the weights and formulas public?

Yes. Component weights, normalization rules, eligibility criteria, and formula versions are documented on the methodology pages and versioned over time. Dated index snapshots preserve past outputs, so scoring changes can be audited instead of silently altering published rankings.

Can I reproduce a score myself?

Yes. The reproducibility documentation and bounded exports provide the normalized inputs, formulas, and source policies needed to recompute published scores, and each model's audit drawer shows exactly which evidence contributed to its score and by how much.