A leaderboard is only as fair as the tests it happens to include.
Simple benchmark averages treat every test as equally informative and independent. That makes rankings sensitive to test selection, overlap, and measurement noise.
The General Intelligence Index combines performance across 50+ benchmarks to estimate a model's latent g factor—not weighted averages—and expresses it on a familiar Artificial IQ (AIQ) scale with a mean of 100 and standard deviation of 15.
Simple benchmark averages treat every test as equally informative and independent. That makes rankings sensitive to test selection, overlap, and measurement noise.
GII weights evidence by difficulty, reliability, discrimination, and shared variance to recover the latent capability that best explains performance across subtests.
GII is a statistically rigorous adaptation of the Epoch Capabilities Index. It models each model’s observed subtest scores as measurements of an unobserved general factor, \(g\).
Under a Gaussian factor model, the covariance structure and posterior score can be written as:
The inverse covariance term \(\boldsymbol{\Sigma}^{-1}\) discounts redundancy. Estimation uses multidimensional item-response theory.
Point estimates are ordered by AIQ. Whiskers show the 90% confidence interval.
| Rank | Model | AIQ | 90% confidence interval |
|---|---|---|---|
| Reading the index… | |||
AIQ means Artificial IQ. These are normalized AI index scores, not claims about consciousness or human equivalence. Confidence intervals may be asymmetric around the point estimate.
AIQ places the latent capability estimate on a scale centered at 100, with a standard deviation of 15. It is a statistical normalization for AI systems—not a claim that model and human intelligence are equivalent.