Language models that do well on one benchmark tend to do well on the others, and published psychometric analyses have extracted a dominant common factor from those scores. Whether that factor is anything like human g depends on a question about what is varying between the models.

Line up a dozen large language models against a dozen benchmarks and the resulting table has an obvious structure. The models near the top of one column are near the top of most of the others. Reading comprehension, grade-school mathematics, code generation, commonsense inference — the rankings rhyme. In human testing that pattern has a name and a century of theory behind it, and the temptation to reach for the same theory here is very strong. It is worth being precise about which part of the analogy holds.
Charles Spearman noticed in 1904 that every pair of mental tests he examined correlated positively — the positive manifold. Extract the shared variance and you get a general factor, g, which in human samples typically accounts for a large share of the differences between people across a diverse battery. Our explainer on what the g factor actually is sets out how it is derived and what it does not mean. The important structural point for what follows: g is estimated from variation between people who each arrived at the test by a separate developmental route.
Applying that machinery to model evaluations is straightforward arithmetic. Take a matrix of models by benchmarks, correlate the columns, run a factor analysis, and see whether one factor dominates. Published work has done exactly this. Ilić and Gignac reported in Intelligence in 2024 that scores across a range of language-model benchmarks are substantially intercorrelated and yield a dominant general factor, and framed the open question as whether that reflects something like general capability or something closer to a shared achievement effect. The finding of a dominant factor is not really in dispute. What it licenses you to say is.
In a human sample, the thing generating variation is roughly a thousand independent causes acting on each person — genetics, schooling, health, environment, motivation on the day. A common factor across tasks under those conditions is informative because there was no obvious single dial that all the differences could have come from. On a model leaderboard, there is a very obvious dial:
A single factor running through every benchmark is what you would expect if the systems differed mainly in how much was spent on them. That is a real finding. It is not the finding people report.
None of those alternatives require the models to share a general cognitive ability. They only require that the models differ along a dominant axis and that most benchmarks are sensitive to it. That is enough to produce a positive manifold and a dominant factor, and it is the null hypothesis a strong claim would have to beat.
It is worth walking the alternative through concretely, because stated abstractly it sounds like pedantry and worked through it does not. Suppose ten models differ in nothing but training compute, spanning two orders of magnitude, and suppose every benchmark score is some increasing function of compute plus noise specific to that benchmark. Correlate the benchmarks across those ten models and every pair comes out positive. Factor-analyse the matrix and one factor dominates. You will have recovered a beautiful positive manifold from a population in which exactly one thing varied, and it was not an ability.
Human samples are not built that way, which is the entire reason the human result carries weight. Two people of similar background and schooling still differ across subtests, and g is estimated from variation that no single input controls. Recovering a dominant factor there is informative because the obvious confound is absent. Recovering one from a leaderboard is uninformative until somebody shows the obvious confound has been removed — and published leaderboard analyses generally cannot, because the models were not built as an experiment.
There is a second problem, and it makes the common factor easier to produce rather than harder. Benchmark questions and their answers are published on the web, and the web is the training corpus. To the extent that test items have been memorised rather than solved, scores rise on every benchmark that has leaked, in every model trained on the same data. That is a correlated inflation across tasks and across models — precisely the pattern a factor analysis reads as evidence of a shared underlying ability.
The question is answerable in principle, and some of the field is already pointed at it. Several kinds of evidence would separate a genuine capability factor from a compute factor:
Until more of that exists, the defensible statement is narrow: benchmark scores across current language models are strongly intercorrelated, a dominant common factor can be extracted from them, and the leading explanations for that factor include several that have nothing to do with general intelligence. Reporting the extraction and skipping the explanations is how a technical result becomes a headline it does not support.
It matters because the language travels. Once a model is described as having an IQ or a general intelligence factor, the number gets compared with human scores, and human scores carry a meaning — a percentile against a norming sample — that the model number does not have. The comparison is not merely imprecise; the two figures are not measurements of the same kind of thing. If you want to know what a properly normed score actually is and what it can bear, our explainer on how IQ tests are built and normed is the place to start, and you can take a properly structured assessment to see the machinery from the inside.
A score means something when you know the sample it was normed against and the error it carries. Take a properly structured assessment and read the percentile rather than the point.
Find your IQ score now! →The honest position is that this is an interesting open question being reported as a settled one. A dominant factor in model benchmark scores is a real, replicable, arithmetically sound finding. Whether it is a fact about the models' abilities or a fact about how much compute went into them is exactly the question, and it is not one that another factor analysis of the same leaderboard can answer. That takes experiments where the confound is controlled rather than averaged over — which is, as it happens, the same standard the intelligence literature had to reach for its own general factor, and it took decades.
Benchmark scores across language models are strongly positively correlated, and published factor analyses — including work by Ilic and Gignac in Intelligence in 2024 — have extracted a dominant common factor from them. Whether that factor represents a general cognitive ability is disputed, because the models on a leaderboard also vary along training compute and overlapping data, which would produce the same pattern.
A human IQ score is a position within a norming sample of people who had not seen the items, reported with a published standard error. A model benchmark score has no norming sample, no error term of that kind, and no guarantee the items were absent from training data. The two numbers are not measurements of the same kind of thing, so a direct comparison has no defined meaning.
It is the presence of benchmark questions and answers in a model's training data, usually because the benchmark was published on the web that the model was trained on. Contaminated items are recalled rather than solved, which inflates scores across every affected benchmark and every model trained on the same corpus — producing correlated score increases with no underlying ability change.
It is built to reduce them. The ARC-AGI series is designed around tasks that are novel to the system at test time, so memorisation of published answers helps less than it does on a standard benchmark. That design is why model performance on it has historically looked very different from performance on conventional benchmarks, and why it is a more informative test of generalisation.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
OpenAI and Anthropic published figures for the same rival models within four days of each other. They do not match, and the reasons are the ones psychometrics named a century ago.
IQ scales are defined by a human reference sample, and a language model is not in it. Placing a machine on that scale is not a hard measurement problem — it is a category error.

Headlines call it an AI intelligence score. It is a weighted average of ten separate evaluations, and the hardest one — an exam built to resist quick solving — is already close to 60% solved.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →