AI Model Index
AI Benchmark Families
Browse AI benchmark families, variants, source policies, evaluation configurations, and the evidence used in composite model indexes.
Frequently asked questions
What is a benchmark family?
A benchmark family groups related evaluations that measure a similar capability, such as coding or reasoning, while keeping individual variants distinct. Different versions, harnesses, agent setups, and scoring policies are never silently merged, so you can see exactly which configuration produced each row.
Why are some benchmark rows excluded from rankings?
Public rankings use only reviewed variants with a declared production adapter. Fixture, demo, seed, unavailable, and unverified historical rows are excluded, and each family page documents its source policy so you can check which rows are eligible and why others are held back.
How do benchmarks connect to the composite indexes?
Normalized rows from eligible benchmark variants feed the published index recipes. The benchmark pages document the source policies and evaluation configurations behind those rows, and every index score traces back to the specific benchmark evidence that contributed to it.