Every properly built test publishes a standard error of measurement alongside the score. On a mean-100, standard-deviation-15 scale it is large enough that a reported 104 and a reported 111 are not reliably different results.
An IQ result arrives as one number, printed once, with no hint that it could have come out differently. That presentation is the problem. Every reputable test manual reports a second figure next to the score that describes how much the score itself wobbles, and once you know that figure most of the arguments people have about IQ points stop being arguments about anything.
A modern IQ score is a position, not a quantity. The test is given to a large reference sample, the results are arranged so that the average sits at 100, and the spread is fixed by a chosen standard deviation — a measure of how far apart people's results fall. A score of 115 on a standard-deviation-15 scale means "one standard deviation above the average of that reference sample", and nothing else. It is not a count of anything.
Because it is a position estimated from a finite set of questions on a particular morning, it inherits every source of noise in that process. The specific items sampled, whether the room was quiet, how well the person slept, whether they had sat something similar before, ordinary fluctuation in attention — all of it moves the number without moving the thing the number is supposed to represent.
Test developers do not leave that noise undescribed. They quantify it as the standard error of measurement, usually shortened to SEM: an estimate of the typical distance between a person's reported score and the score they would average over many independent sittings. It is derived from the test's reliability and its standard deviation, and it is published in the technical manual for every score the test produces.
For a full-scale score on a well-constructed adult battery scored with a standard deviation of 15, that error is commonly in the region of 2 to 3 points. Turning it into a confidence interval is ordinary arithmetic — about 1.96 standard errors either side for a 95 per cent interval:
The reported figure is the middle of a band the test itself tells you how to draw. Quoting the middle and discarding the band is not precision, it is omission.
Set two results side by side. One person reports 104, another reports 111. The gap looks real — seven points, and one of them is above average by more. But each of those numbers carries a band of roughly five points either side, and the bands overlap across almost their whole width. There is no basis in the scores for saying the second person scored higher in any durable sense. The honest statement is that the two results are not distinguishable.
The same logic applies to one person tested twice. A rise from 104 to 111 between two sittings is comfortably inside what measurement error alone produces, before considering the practice effect — the well-documented tendency for a second sitting on a familiar format to come out higher regardless of any change in ability. Our piece on what a retest actually measures works through that separate problem in detail.
The full-scale score is the most reliable number a battery produces, because it pools the most items. Everything beneath it is shorter and therefore noisier, and the published standard errors reflect that:
None of that makes the scores useless. It makes them estimates with known precision, which is a great deal better than most numbers people quote about themselves. The failure is not in the test, it is in reporting one figure from it as though it were exact. The same question of build quality applies before any of this arithmetic is worth doing — an unnormed test has no meaningful error term at all, because there is no reference sample for the score to be an estimate of.
Three things, in order. Name the scale it was measured on, because a bare number is uninterpretable without its mean and standard deviation. Convert it to a percentile against that scale, which is the form that survives translation between tests. Then read it as a range rather than a point, and describe it that way to anyone who asks. Our IQ percentile calculator does the first two conversions, and treating the answer as a band rather than a figure is the whole discipline.
A score is worth having when you know what it is an estimate of and how precise that estimate is. Take a properly normed assessment, read the percentile rather than the point, and keep the confidence range attached to it.
Find your IQ score now! →The habit generalises well beyond testing. Any measurement of a person made once, on one day, from a sample of their behaviour, is an estimate with a band around it. Test manuals are unusual mainly in publishing the band. When someone quotes a score without one — their own, a public figure's, a country's average — the missing interval is not a detail that was left out for brevity. It is the part that tells you whether the number can bear the weight being put on it.
It is published as the standard error of measurement in the test's technical manual. For a full-scale score on a well-normed adult battery using a mean of 100 and a standard deviation of 15, it is commonly in the region of 2 to 3 points, which gives a 95 per cent confidence interval of roughly 5 to 6 points either side of the reported figure. Index and subtest scores carry larger errors than the full-scale total.
Usually not, on its own. If each score carries a confidence interval of about 5 points either side, two results seven points apart produce bands that overlap across most of their width, and the scores are not reliably distinguishable. That applies both to comparing two people and to comparing one person's results across two sittings.
Two ordinary reasons before any change in ability is considered. Measurement error alone moves a result by several points between sittings, and the practice effect tends to raise a second score on a familiar format. A retest that lands inside the first score's confidence interval is the expected outcome, not evidence of a change.
No. It means the test reports its own precision, which is a mark of a well-built instrument rather than a flaw. A test that publishes no standard error is not more accurate; it is simply not telling you how accurate it is.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
Wechsler scales use a standard deviation of 15, older Stanford-Binet forms used 16, and the Cattell scales use 24. One standing prints as three different numbers — and one printed 140 is a one-in-260 result on one of those scales and a one-in-21 result on another.
On the scale most modern tests use, 120 sits around the 91st percentile. Change the scale and the same number moves. Add the measurement error every test carries and it stops being a point at all.
Mensa admits at the 98th percentile of a supervised, properly normed test. That single rule produces a different qualifying figure on every scale — and none of them is the 130 the internet keeps repeating.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →