IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

An IQ Score Is a Range, Not a Number: The Margin of Error

Every properly built test publishes a standard error of measurement alongside the score. On a mean-100, standard-deviation-15 scale it is large enough that a reported 104 and a reported 111 are not reliably different results.

Illustration generated for IQ Metrics. No photograph is used. IQ Metrics

What you need to know

  • A reported IQ is a point estimate, not a measurement of a fixed quantity. Test manuals publish a standard error of measurement beside it, describing how far a result moves between sittings for reasons that have nothing to do with the person.
  • On well-normed adult scales with a mean of 100 and a standard deviation of 15, the standard error for a full-scale score is commonly in the region of 2 to 3 points. A 95 per cent confidence interval is therefore roughly 5 to 6 points either side of the reported figure.
  • That interval is why a reported 104 and a reported 111 are not reliably different: both are consistent with the same underlying standing. Comparing two people, or the same person across two sittings, means comparing intervals rather than numbers.
  • The interval is wider for index and subtest scores than for the full-scale total, because they rest on fewer items. A single subtest figure quoted on its own is the least reliable number the test produces.

An IQ result arrives as one number, printed once, with no hint that it could have come out differently. That presentation is the problem. Every reputable test manual reports a second figure next to the score that describes how much the score itself wobbles, and once you know that figure most of the arguments people have about IQ points stop being arguments about anything.

What the reported number actually is

A modern IQ score is a position, not a quantity. The test is given to a large reference sample, the results are arranged so that the average sits at 100, and the spread is fixed by a chosen standard deviation — a measure of how far apart people's results fall. A score of 115 on a standard-deviation-15 scale means "one standard deviation above the average of that reference sample", and nothing else. It is not a count of anything.

Because it is a position estimated from a finite set of questions on a particular morning, it inherits every source of noise in that process. The specific items sampled, whether the room was quiet, how well the person slept, whether they had sat something similar before, ordinary fluctuation in attention — all of it moves the number without moving the thing the number is supposed to represent.

The standard error of measurement

Test developers do not leave that noise undescribed. They quantify it as the standard error of measurement, usually shortened to SEM: an estimate of the typical distance between a person's reported score and the score they would average over many independent sittings. It is derived from the test's reliability and its standard deviation, and it is published in the technical manual for every score the test produces.

For a full-scale score on a well-constructed adult battery scored with a standard deviation of 15, that error is commonly in the region of 2 to 3 points. Turning it into a confidence interval is ordinary arithmetic — about 1.96 standard errors either side for a 95 per cent interval:

  • A standard error of 2.5 points gives a 95 per cent interval of roughly 5 points either side.
  • A reported 112 therefore describes a person whose true standing is plausibly anywhere from about 107 to about 117.
  • A reported 100 describes a range of roughly 95 to 105 — which spans, on an SD-15 scale, something like the 37th to the 63rd percentile.
  • The interval belongs to the score, not to the person. Repeating the test does not narrow it; it simply produces a second number inside the same band.

The reported figure is the middle of a band the test itself tells you how to draw. Quoting the middle and discarding the band is not precision, it is omission.

Why two different scores can mean the same thing

Set two results side by side. One person reports 104, another reports 111. The gap looks real — seven points, and one of them is above average by more. But each of those numbers carries a band of roughly five points either side, and the bands overlap across almost their whole width. There is no basis in the scores for saying the second person scored higher in any durable sense. The honest statement is that the two results are not distinguishable.

The same logic applies to one person tested twice. A rise from 104 to 111 between two sittings is comfortably inside what measurement error alone produces, before considering the practice effect — the well-documented tendency for a second sitting on a familiar format to come out higher regardless of any change in ability. Our piece on what a retest actually measures works through that separate problem in detail.

This is also why threshold claims need care. A society or programme that admits at a fixed cut score is drawing an administrative line through a continuum, not separating two kinds of person: a reported score a point below the line and one a point above it are the same result as far as the test can tell. That is the arithmetic behind our note on what the Mensa requirement really is, where the qualifying figure changes with the scale while the percentile behind it does not.

Where the interval gets wider

The full-scale score is the most reliable number a battery produces, because it pools the most items. Everything beneath it is shorter and therefore noisier, and the published standard errors reflect that:

  • Index scores — the groupings for verbal comprehension, working memory, processing speed and so on — typically carry larger standard errors than the full-scale total, so their intervals are wider.
  • Individual subtest scores are shorter still and are the least stable figures on the report. A subtest result quoted on its own carries very little information.
  • Differences between two index scores inherit the error of both, so a gap has to be substantial before it means anything. Manuals publish tables of how large a difference must be to be treated as real.
  • Short-form and abbreviated tests trade items for time, and their intervals are wider than the full battery they approximate.

None of that makes the scores useless. It makes them estimates with known precision, which is a great deal better than most numbers people quote about themselves. The failure is not in the test, it is in reporting one figure from it as though it were exact. The same question of build quality applies before any of this arithmetic is worth doing — an unnormed test has no meaningful error term at all, because there is no reference sample for the score to be an estimate of.

What to do with a score you already have

Three things, in order. Name the scale it was measured on, because a bare number is uninterpretable without its mean and standard deviation. Convert it to a percentile against that scale, which is the form that survives translation between tests. Then read it as a range rather than a point, and describe it that way to anyone who asks. Our IQ percentile calculator does the first two conversions, and treating the answer as a band rather than a figure is the whole discipline.

Your own number

Where would your own score land?

A score is worth having when you know what it is an estimate of and how precise that estimate is. Take a properly normed assessment, read the percentile rather than the point, and keep the confidence range attached to it.

Find your IQ score now!
Secure & encryptedInstant results10–20 minutes

The habit generalises well beyond testing. Any measurement of a person made once, on one day, from a sample of their behaviour, is an estimate with a band around it. Test manuals are unusual mainly in publishing the band. When someone quotes a score without one — their own, a public figure's, a country's average — the missing interval is not a detail that was left out for brevity. It is the part that tells you whether the number can bear the weight being put on it.

Common questions

What is the margin of error on an IQ test?

It is published as the standard error of measurement in the test's technical manual. For a full-scale score on a well-normed adult battery using a mean of 100 and a standard deviation of 15, it is commonly in the region of 2 to 3 points, which gives a 95 per cent confidence interval of roughly 5 to 6 points either side of the reported figure. Index and subtest scores carry larger errors than the full-scale total.

Is a 7-point difference in IQ scores meaningful?

Usually not, on its own. If each score carries a confidence interval of about 5 points either side, two results seven points apart produce bands that overlap across most of their width, and the scores are not reliably distinguishable. That applies both to comparing two people and to comparing one person's results across two sittings.

Why did my IQ score change when I took the test again?

Two ordinary reasons before any change in ability is considered. Measurement error alone moves a result by several points between sittings, and the practice effect tends to raise a second score on a familiar format. A retest that lands inside the first score's confidence interval is the expected outcome, not evidence of a change.

Does a confidence interval mean the test is unreliable?

No. It means the test reports its own precision, which is a mark of a well-built instrument rather than a flaw. A test that publishes no standard error is not more accurate; it is simply not telling you how accurate it is.

Sources for this story

  1. Technical and interpretive manuals for the Wechsler intelligence scales, on reliability, the standard error of measurement and confidence intervals — Pearson
  2. Standards for Educational and Psychological Testing, on reliability, precision and the reporting of measurement error — American Educational Research Association, American Psychological Association and National Council on Measurement in Education
  3. Stanford-Binet Intelligence Scales technical manual, on score reliability and difference-score tables — Riverside Insights
  4. Psychometric Theory, on classical test theory and the derivation of the standard error of measurement — Nunnally and Bernstein, McGraw-Hill
  5. International Guidelines for Test Use, on reporting scores as ranges rather than point values — International Test Commission

Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.

Share this story

Know someone who keeps seeing this number quoted without the scale it was measured on? Send it to them — it takes one tap.

Filed under#confidence intervals#scores and scales#percentiles and norms#wechsler scales#test conditions

Read the research.
Then find your own number.

Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.

Start IQ Test
Secure & encryptedInstant results10–20 minutes