Subtracting one IQ score from another is not a comparison. Both numbers carry a margin of error, and most of the gaps people argue about disappear inside it. Put two scores in and get a straight answer.
How many people are in the gap
Ten points near the middle separates a quarter of everyone. The same ten points out at the edge separates almost nobody. Hover to read any point.
Not by eye, and not by subtraction. A difference has to be bigger than the noise in the two measurements before it counts as a difference at all.
SEM = SD × √(1 − r)
No test measures perfectly. The standard error of measurement is how far a score typically drifts if the same person sits the same test again. At SD 15 and r = .90 that is 4.7 points.
SEdiff = SEM × √2
You are comparing two imperfect numbers, so both errors are in play. That is 6.7 points of noise in the gap itself before anything real has happened.
critical = 1.96 × SEdiff
The gap must clear this to be called real with 95% confidence. At SD 15 and r = .90 that is 13.1 points — which is why a 10-point gap does not survive.
It is tempting to look at the two ranges above and say “they overlap, so the difference is not real.” That test is wrong. Two 95% confidence intervals can overlap substantially while the difference between them is still statistically significant — judging by overlap is conservative in the wrong direction and rejects real differences. The bars are drawn so you can see where each score sits; the verdict is decided by the critical difference, which is the correct test.
Even when a gap is real, its size in points tells you almost nothing. Because the curve is not flat, an identical ten-point gap covers wildly different shares of the population depending on where it sits.
26.1%
Roughly one person in four sits inside this gap. Ten points here separates a large slice of the population.
11.1%
The same ten points, less than half the population between them. The curve has already started to thin out.
1.9%
Still ten points, now fewer than two people in a hundred. Thirteen times less population than the same gap in the middle.
This is why “I scored ten points higher” is not a statement about ability until you say where. The bars in the tool above are drawn on the curve for exactly this reason — position carries meaning that a subtraction throws away.
The maths above assumes both scores came from the same kind of instrument, taken under comparable conditions. When that is not true, no amount of arithmetic makes the comparison valid.
A 148 from Cattell III B and a 130 from a Wechsler test are the same result, not an 18-point gap. Convert both to one scale before comparing anything — the score converter does this.
Even on the same scale, a supervised clinical battery and a ten-minute online quiz are not measuring with the same instrument. The critical difference above assumes one instrument used twice; across different tests the true uncertainty is larger, sometimes much larger.
Extreme scores contain more luck than middling ones, so a retest tends to land nearer the average. That drift is regression to the mean, not a change in ability, and it is the most commonly misread pattern in repeat testing.
Usually not. On the SD 15 scale, two scores from a test with reliability r = .90 need to differ by about 13 points before the difference clears the margin of error at 95% confidence. A 10-point gap falls inside the noise. On a more reliable test — r = .95 — the threshold drops to about 9 points and the same 10-point gap does become significant. The reliability of the test matters more than the size of the gap.
Compare the gap against the critical difference, which is 1.96 × SEM × √2, where SEM = SD × √(1 − reliability). If the gap is larger, the difference is real at 95% confidence. If it is smaller, the two scores are not distinguishable and the same person could have produced either one on a different day. The tool above does this calculation for you.
Only with caution, and only after converting both to the same scale. Different tests measure somewhat different things and have different reliabilities, so the real uncertainty in the comparison is larger than the figure this tool reports. Comparing a supervised clinical battery with a short online test is not a meaningful comparison at all, regardless of the numbers.
Because the overlap test is the wrong test. Two 95% confidence intervals can overlap noticeably while the difference between the scores is still statistically significant. The correct comparison uses the standard error of the difference, which is smaller than simply laying the two intervals side by side would suggest. Judging by overlap rejects real differences.
Probably not. Two things are usually at work: the normal measurement error, which moves scores several points in either direction with no change in ability, and regression to the mean, which pulls unusually high or low results back toward the average on retesting. Check the gap against the critical difference before treating a change as real.
Yes. It runs entirely in your browser, is free to use and needs no account. Nothing you type is sent anywhere.
A comparison is only as good as the scores going into it. Take an assessment that reports its scale, its percentile and its confidence range instead of a bare number.