A latent measure of AI capability

One signal for
general intelligence.

The General Intelligence Index combines performance across 50+ benchmarks to estimate a model's latent g factor—not weighted averages—and expresses it on a familiar Artificial IQ (AIQ) scale with a mean of 100 and standard deviation of 15.

01 The problem

A leaderboard is only as fair as the tests it happens to include.

Simple benchmark averages treat every test as equally informative and independent. That makes rankings sensitive to test selection, overlap, and measurement noise.

02 The solution

Measure the ability shared across many different kinds of tasks.

GII weights evidence by difficulty, reliability, discrimination, and shared variance to recover the latent capability that best explains performance across subtests.

Statistical foundation How is the index calculated? Expand

GII is a statistically rigorous adaptation of the Epoch Capabilities Index. It models each model’s observed subtest scores as measurements of an unobserved general factor, \(g\).

\[ \mathbf{x} = \boldsymbol{\nu} + \boldsymbol{\Lambda}g + \boldsymbol{\epsilon} \] \(\boldsymbol{\Lambda}\) = \(g\)-loadings  ·  \(\boldsymbol{\epsilon}\) = residual performance

Why is this better than a weighted average?

  • A reliable, highly \(g\)-loaded subtest carries more information.
  • A weakly \(g\)-loaded subtest has less influence on the estimate.
  • Strongly correlated subtests are treated as partly redundant—not as two independent wins.
  • Measurement error, discrimination, and item difficulty can be incorporated explicitly.

Under a Gaussian factor model, the covariance structure and posterior score can be written as:

\[ \boldsymbol{\Sigma} = \boldsymbol{\Lambda}\Phi\boldsymbol{\Lambda}^{\mathsf T} + \boldsymbol{\Theta} \] \[ \mathbb{E}\!\left[g \mid \mathbf{x}\right] = \Phi\boldsymbol{\Lambda}^{\mathsf T}\boldsymbol{\Sigma}^{-1} \left(\mathbf{x}-\boldsymbol{\nu}\right) \]

The inverse covariance term \(\boldsymbol{\Sigma}^{-1}\) discounts redundancy. Estimation uses multidimensional item-response theory.

Model intelligence

The index, ranked.

Point estimates are ordered by AIQ. Whiskers show the 90% confidence interval.

Loading models… 90% confidence interval
Rank Model AIQ 90% confidence interval
Reading the index…

AIQ means Artificial IQ. These are normalized AI index scores, not claims about consciousness or human equivalence. Confidence intervals may be asymmetric around the point estimate.

The normalized scale

AIQ Scores

AIQ places the latent capability estimate on a scale centered at 100, with a standard deviation of 15. It is a statistical normalization for AI systems—not a claim that model and human intelligence are equivalent.

AIQ bell curve from 55 to 145 A normal distribution centered at AIQ 100, with markings every 15 points from 55 through 145. 55−3σ 70−2σ 85−1σ 100mean 115+1σ 130+2σ 145+3σ
Artificial IQ (AIQ) · mean 100 · standard deviation 15