IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

When a Model “Scores 120”, That Number Is Not an IQ

IQ scales are defined by a human reference sample, and a language model is not in it. Placing a machine on that scale is not a hard measurement problem — it is a category error.

Illustration generated for IQ Metrics. No photograph is used. IQ Metrics

What you need to know

  • An IQ score is a position within a reference sample of people. A model is not a member of that sample, so a percentile computed for it describes nothing.
  • Public test items are on the public internet, which means they are plausibly in the training data — a model may be recalling rather than reasoning.
  • IQ tests are validated against human outcomes such as school and job performance. None of that validity evidence transfers to a machine.
  • The benchmarks built specifically to resist memorisation exist precisely because scores on repurposed human tests turned out to be so easy to misread.

Every few months a chart circulates placing the current crop of AI models on an IQ scale, usually with a line for the human average and an arrow going up and to the right. It is a compelling image. It is also measuring something other than what it claims, and the reasons are worth stating precisely, because two of them apply to human testing as well.

A score is a position in a sample

This is the fundamental problem and it is not a technicality. An IQ score is a deviation score. The test is given to a large reference sample chosen to represent a human population; the average of that sample is set to 100 and the spread is set so one standard deviation equals a fixed number of points. Your score says where you fall within that distribution.

A model is not in the reference sample and is not a member of the population it represents. Computing a percentile for it is like computing what percentile a submarine is at in a swimming competition — the arithmetic runs, and the output does not describe anything. This is not a limitation of current models or of current tests. It is what the scale is.

The arithmetic will happily produce a number. That is not the same as the number meaning something.

Contamination

The second problem is more familiar and more fixable. Published test items, sample questions, practice sets and full retired forms are all on the public web, and models are trained on large web corpora. When a model answers a well-known item correctly, there is no straightforward way to tell from the outside whether it reasoned to the answer or has effectively seen it before.

Researchers call this benchmark contamination, and it is a recognised and active problem across machine learning evaluation generally, not something specific to IQ-style tests. The reason it bites hardest here is that IQ item banks are old, widely reproduced and heavily discussed — which is exactly the profile of material that ends up in a training corpus many times over.

Validity does not transfer

The third problem is the one that gets skipped. A test is not valid in the abstract; it is valid for a purpose, in a population, supported by evidence. The evidence supporting IQ tests consists of decades of studies relating human scores to human outcomes — school performance, training success, job performance, and so on. That is what the scores were validated to predict.

None of that evidence says anything about a machine. Even if a model produced a score under perfectly clean conditions, there is no body of work connecting that score to anything a model does. The number would have no interpretation attached, which is a strange thing for a measurement to be missing.

This is why the benchmarks designed for machines look the way they do. Evaluations such as ARC-AGI are built around tasks that are easy to generate freshly and hard to have memorised, with held-out sets kept off the public web, precisely so that a score reflects generalisation rather than recall. Whatever you think of any particular benchmark, the design goal is the right one and repurposed human IQ items do not meet it.

What a legitimate claim would look like

  • Items generated or held out after the model’s training cutoff, with the cutoff stated.
  • The full protocol published — prompting, number of attempts, whether tools were available, how ties and refusals were scored.
  • Results reported as performance on a named task, not as a position on a human scale.
  • An explicit statement of what the score is evidence for, since IQ validity evidence does not carry over.
  • Repeated runs with variation reported, because a single run of a stochastic system is one sample, not a measurement.

Two of those five — state your protocol, report your uncertainty — are just ordinary measurement hygiene, and the fact that AI evaluation and human testing need the same reminders is not a coincidence. Bad measurement fails the same way regardless of what is being measured.

Your own number

Where would your own score land?

The human version of this is worth holding to the same standard. A score means something when it names the scale it sits on, states the percentile against a defined reference sample, and reports the confidence range around it — because a single sitting produces an estimate rather than a fixed property. A number with none of those attached is a headline, whether it belongs to a person or to a model.

Find your IQ score now!
Secure & encryptedInstant results10–20 minutes

Models are getting measurably better at hard reasoning tasks, and that is genuinely interesting and worth reporting carefully. Reporting it as an IQ is the one framing guaranteed to say less than the underlying result does.

Common questions

Can an AI have an IQ?

Not in any meaningful sense. An IQ score is a position within a reference sample of people, and a model is not a member of that sample or of the population it represents. The arithmetic can be performed; the resulting percentile does not describe anything.

Why do published AI IQ scores keep going up?

Because models are improving at the kinds of reasoning tasks these items use, and because widely published items are plausibly present in training data, which makes correct answers ambiguous between reasoning and recall. The rising line reflects both, and the two cannot be separated from the outside.

Is there a fair way to compare human and machine reasoning?

On specific, well-specified tasks, yes — with items held out or generated after the training cutoff, the full protocol published, and results reported as task performance rather than as a position on a human scale. Benchmarks such as ARC-AGI are built around exactly that design goal.

Sources for this story

  1. Standards for Educational and Psychological Testing, on norming samples, reference populations and construct validity — American Educational Research Association, American Psychological Association and National Council on Measurement in Education
  2. ARC-AGI benchmark documentation on held-out task design and resistance to memorisation — ARC Prize Foundation
  3. Research literature on benchmark contamination in large language model evaluation — Machine learning conference proceedings and preprint archives
  4. Technical and interpretive manuals for the Wechsler intelligence scales, covering deviation scoring — Pearson

Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.

Share this story

Know someone who keeps seeing this number quoted without the scale it was measured on? Send it to them — it takes one tap.

Filed under#artificial intelligence#ai benchmarks#percentiles and norms#study quality

Read the research.
Then find your own number.

Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.

Start IQ Test
Secure & encryptedInstant results10–20 minutes