IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

Is There a g Factor for AI Models? What the Benchmark Correlations Show

Language models that do well on one benchmark tend to do well on the others, and published psychometric analyses have extracted a dominant common factor from those scores. Whether that factor is anything like human g depends on a question about what is varying between the models.

Is There a g Factor for AI Models? What the Benchmark Correlations Show
Illustration generated for IQ Metrics. No photograph is used. IQ Metrics

What you need to know

  • Benchmark scores across language models are strongly positively correlated: a model that leads on reading comprehension tends to lead on mathematics, code and commonsense reasoning too. Factor analyses published in the psychometric literature have extracted a single dominant factor from those score matrices.
  • That is formally the same operation Spearman performed on human test scores in 1904, and it is why the result gets reported as "AI has a g factor". The arithmetic is genuine; the interpretation is where the argument is.
  • The critical disanalogy is what varies. Human g is extracted from differences between people who each developed independently. A model leaderboard varies mostly along training compute, data volume and overlapping corpora — so a single common factor is at least as consistent with "these systems differ mainly in scale" as with "these systems share a general ability".
  • Contamination makes it worse rather than better: benchmark items that appear in training data inflate correlated scores across every model trained on the same web, producing a common factor with no cognitive content at all.

Line up a dozen large language models against a dozen benchmarks and the resulting table has an obvious structure. The models near the top of one column are near the top of most of the others. Reading comprehension, grade-school mathematics, code generation, commonsense inference — the rankings rhyme. In human testing that pattern has a name and a century of theory behind it, and the temptation to reach for the same theory here is very strong. It is worth being precise about which part of the analogy holds.

What the pattern is, in human testing

Charles Spearman noticed in 1904 that every pair of mental tests he examined correlated positively — the positive manifold. Extract the shared variance and you get a general factor, g, which in human samples typically accounts for a large share of the differences between people across a diverse battery. Our explainer on what the g factor actually is sets out how it is derived and what it does not mean. The important structural point for what follows: g is estimated from variation between people who each arrived at the test by a separate developmental route.

Applying that machinery to model evaluations is straightforward arithmetic. Take a matrix of models by benchmarks, correlate the columns, run a factor analysis, and see whether one factor dominates. Published work has done exactly this. Ilić and Gignac reported in Intelligence in 2024 that scores across a range of language-model benchmarks are substantially intercorrelated and yield a dominant general factor, and framed the open question as whether that reflects something like general capability or something closer to a shared achievement effect. The finding of a dominant factor is not really in dispute. What it licenses you to say is.

The disanalogy that matters

In a human sample, the thing generating variation is roughly a thousand independent causes acting on each person — genetics, schooling, health, environment, motivation on the day. A common factor across tasks under those conditions is informative because there was no obvious single dial that all the differences could have come from. On a model leaderboard, there is a very obvious dial:

  • Training compute, which varies across orders of magnitude between the models on a typical leaderboard.
  • Training data volume and overlap — most large models are trained on heavily overlapping web corpora, so they are not independent samples in any meaningful sense.
  • Post-training technique: instruction tuning, preference optimisation and reasoning-oriented training raise scores across many benchmarks at once.
  • Evaluation harness and prompting, which our note on why the same model scores differently on the same benchmark shows can move a published figure substantially on its own.

A single factor running through every benchmark is what you would expect if the systems differed mainly in how much was spent on them. That is a real finding. It is not the finding people report.

None of those alternatives require the models to share a general cognitive ability. They only require that the models differ along a dominant axis and that most benchmarks are sensitive to it. That is enough to produce a positive manifold and a dominant factor, and it is the null hypothesis a strong claim would have to beat.

It is worth walking the alternative through concretely, because stated abstractly it sounds like pedantry and worked through it does not. Suppose ten models differ in nothing but training compute, spanning two orders of magnitude, and suppose every benchmark score is some increasing function of compute plus noise specific to that benchmark. Correlate the benchmarks across those ten models and every pair comes out positive. Factor-analyse the matrix and one factor dominates. You will have recovered a beautiful positive manifold from a population in which exactly one thing varied, and it was not an ability.

Human samples are not built that way, which is the entire reason the human result carries weight. Two people of similar background and schooling still differ across subtests, and g is estimated from variation that no single input controls. Recovering a dominant factor there is informative because the obvious confound is absent. Recovering one from a leaderboard is uninformative until somebody shows the obvious confound has been removed — and published leaderboard analyses generally cannot, because the models were not built as an experiment.

Contamination is not a side issue

There is a second problem, and it makes the common factor easier to produce rather than harder. Benchmark questions and their answers are published on the web, and the web is the training corpus. To the extent that test items have been memorised rather than solved, scores rise on every benchmark that has leaked, in every model trained on the same data. That is a correlated inflation across tasks and across models — precisely the pattern a factor analysis reads as evidence of a shared underlying ability.

This is why an "AI IQ score" is a category error rather than merely an imprecise figure. A human IQ score means something because it is a position in a norming sample of people who did not see the questions in advance, on a test whose reliability and standard error are published. None of those conditions hold for a model on a public benchmark. Our piece on why a model scoring on an IQ-style test has not taken an IQ test works through what breaks, and it is most of the machinery.

What would count as better evidence

The question is answerable in principle, and some of the field is already pointed at it. Several kinds of evidence would separate a genuine capability factor from a compute factor:

  • Holding compute and training data roughly constant and varying architecture or method, then asking whether a common factor survives. If it is a scale artefact, it should collapse.
  • Reporting results at the level of individual test items rather than benchmark averages — the case made by Burnell and colleagues in Science in 2023 — which lets you see whether failures cluster the way an ability model predicts.
  • Benchmarks built to resist memorisation, which is the explicit design goal of the ARC-AGI series; our coverage of how models perform on ARC-AGI-3 shows how differently the picture looks on a test constructed for novelty.
  • Task batteries designed from a theory of what is being measured, rather than assembled from whatever benchmarks were available, which is the direction Hernández-Orallo and colleagues have argued for.

Until more of that exists, the defensible statement is narrow: benchmark scores across current language models are strongly intercorrelated, a dominant common factor can be extracted from them, and the leading explanations for that factor include several that have nothing to do with general intelligence. Reporting the extraction and skipping the explanations is how a technical result becomes a headline it does not support.

Why this matters outside AI research

It matters because the language travels. Once a model is described as having an IQ or a general intelligence factor, the number gets compared with human scores, and human scores carry a meaning — a percentile against a norming sample — that the model number does not have. The comparison is not merely imprecise; the two figures are not measurements of the same kind of thing. If you want to know what a properly normed score actually is and what it can bear, our explainer on how IQ tests are built and normed is the place to start, and you can take a properly structured assessment to see the machinery from the inside.

Your own number

Where would your own score land?

A score means something when you know the sample it was normed against and the error it carries. Take a properly structured assessment and read the percentile rather than the point.

Find your IQ score now!
Secure & encryptedInstant results10–20 minutes

The honest position is that this is an interesting open question being reported as a settled one. A dominant factor in model benchmark scores is a real, replicable, arithmetically sound finding. Whether it is a fact about the models' abilities or a fact about how much compute went into them is exactly the question, and it is not one that another factor analysis of the same leaderboard can answer. That takes experiments where the confound is controlled rather than averaged over — which is, as it happens, the same standard the intelligence literature had to reach for its own general factor, and it took decades.

Common questions

Do AI models have a g factor?

Benchmark scores across language models are strongly positively correlated, and published factor analyses — including work by Ilic and Gignac in Intelligence in 2024 — have extracted a dominant common factor from them. Whether that factor represents a general cognitive ability is disputed, because the models on a leaderboard also vary along training compute and overlapping data, which would produce the same pattern.

Why is comparing an AI benchmark score to a human IQ misleading?

A human IQ score is a position within a norming sample of people who had not seen the items, reported with a published standard error. A model benchmark score has no norming sample, no error term of that kind, and no guarantee the items were absent from training data. The two numbers are not measurements of the same kind of thing, so a direct comparison has no defined meaning.

What is benchmark contamination?

It is the presence of benchmark questions and answers in a model's training data, usually because the benchmark was published on the web that the model was trained on. Contaminated items are recalled rather than solved, which inflates scores across every affected benchmark and every model trained on the same corpus — producing correlated score increases with no underlying ability change.

Does ARC-AGI avoid these problems?

It is built to reduce them. The ARC-AGI series is designed around tasks that are novel to the system at test time, so memorisation of published answers helps less than it does on a standard benchmark. That design is why model performance on it has historically looked very different from performance on conventional benchmarks, and why it is a more informative test of generalisation.

Sources for this story

  1. Evidence of interrelated cognitive-like capabilities in large language models, on the extraction of a general factor from benchmark scores — Ilic and Gignac, Intelligence, 2024
  2. Rethink reporting of evaluation results in AI, on instance-level rather than aggregate reporting — Burnell et al., Science, 2023
  3. On the Measure of Intelligence, on generalisation, priors and the design of the ARC tasks — Chollet, 2019
  4. The Measure of All Minds: Evaluating Natural and Artificial Intelligence, on task batteries built from a measurement theory — Hernandez-Orallo, Cambridge University Press, 2017
  5. General intelligence objectively determined and measured, the original account of the positive manifold — Spearman, American Journal of Psychology, 1904

Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.

Share this story

Know someone who keeps seeing this number quoted without the scale it was measured on? Send it to them — it takes one tap.

Filed under#ai benchmarks#fluid reasoning#study quality#scores and scales

Read the research.
Then find your own number.

Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.

Start IQ Test
Secure & encryptedInstant results10–20 minutes