IQ scales are defined by a human reference sample, and a language model is not in it. Placing a machine on that scale is not a hard measurement problem — it is a category error.
Every few months a chart circulates placing the current crop of AI models on an IQ scale, usually with a line for the human average and an arrow going up and to the right. It is a compelling image. It is also measuring something other than what it claims, and the reasons are worth stating precisely, because two of them apply to human testing as well.
This is the fundamental problem and it is not a technicality. An IQ score is a deviation score. The test is given to a large reference sample chosen to represent a human population; the average of that sample is set to 100 and the spread is set so one standard deviation equals a fixed number of points. Your score says where you fall within that distribution.
A model is not in the reference sample and is not a member of the population it represents. Computing a percentile for it is like computing what percentile a submarine is at in a swimming competition — the arithmetic runs, and the output does not describe anything. This is not a limitation of current models or of current tests. It is what the scale is.
The arithmetic will happily produce a number. That is not the same as the number meaning something.
The second problem is more familiar and more fixable. Published test items, sample questions, practice sets and full retired forms are all on the public web, and models are trained on large web corpora. When a model answers a well-known item correctly, there is no straightforward way to tell from the outside whether it reasoned to the answer or has effectively seen it before.
Researchers call this benchmark contamination, and it is a recognised and active problem across machine learning evaluation generally, not something specific to IQ-style tests. The reason it bites hardest here is that IQ item banks are old, widely reproduced and heavily discussed — which is exactly the profile of material that ends up in a training corpus many times over.
The third problem is the one that gets skipped. A test is not valid in the abstract; it is valid for a purpose, in a population, supported by evidence. The evidence supporting IQ tests consists of decades of studies relating human scores to human outcomes — school performance, training success, job performance, and so on. That is what the scores were validated to predict.
None of that evidence says anything about a machine. Even if a model produced a score under perfectly clean conditions, there is no body of work connecting that score to anything a model does. The number would have no interpretation attached, which is a strange thing for a measurement to be missing.
Two of those five — state your protocol, report your uncertainty — are just ordinary measurement hygiene, and the fact that AI evaluation and human testing need the same reminders is not a coincidence. Bad measurement fails the same way regardless of what is being measured.
The human version of this is worth holding to the same standard. A score means something when it names the scale it sits on, states the percentile against a defined reference sample, and reports the confidence range around it — because a single sitting produces an estimate rather than a fixed property. A number with none of those attached is a headline, whether it belongs to a person or to a model.
Find your IQ score now! →Models are getting measurably better at hard reasoning tasks, and that is genuinely interesting and worth reporting carefully. Reporting it as an IQ is the one framing guaranteed to say less than the underlying result does.
Not in any meaningful sense. An IQ score is a position within a reference sample of people, and a model is not a member of that sample or of the population it represents. The arithmetic can be performed; the resulting percentile does not describe anything.
Because models are improving at the kinds of reasoning tasks these items use, and because widely published items are plausibly present in training data, which makes correct answers ambiguous between reasoning and recall. The rising line reflects both, and the two cannot be separated from the outside.
On specific, well-specified tasks, yes — with items held out or generated after the training cutoff, the full protocol published, and results reported as task performance rather than as a position on a human scale. Benchmarks such as ARC-AGI are built around exactly that design goal.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
The sturdiest results are narrow and decades old. The 2025 studies that drew the headlines rest on self-reports and small preprints — and none of them measured intelligence at all.
On the scale most modern tests use, 120 sits around the 91st percentile. Change the scale and the same number moves. Add the measurement error every test carries and it stops being a point at all.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →