A score of 160 is four standard deviations above average — about one person in 31,560. The sample the score is measured against holds roughly 2,200 people, which is why the top of the table is calculated rather than observed.
Numbers like 200, 230 and 250 circulate freely as the IQs of famous people, and they share a property that is easy to miss: no mainstream individually administered intelligence test can produce any of them. The ceiling on the tests a psychologist actually uses sits near 160, and it sits there for reasons that are arithmetic rather than policy. Understanding why is the fastest route to understanding what a deviation IQ is in the first place.
A modern IQ is a deviation score. The test does not measure an amount of intelligence the way a scale measures an amount of flour. It measures where a person's performance falls within a reference sample — the standardisation sample — and then prints that position on a scale where the average is set to 100 and the spread is set by a chosen standard deviation, usually 15.
Everything follows from that. If the score is a position within a sample, the highest reportable score is limited by the sample. You cannot report a position finer than the sample can resolve, and a sample of a few thousand people cannot resolve positions that occur once in tens of thousands.
The major Wechsler scales report a full-scale IQ across a range of roughly 40 to 160, and their standardisation samples run to about 2,200 people. Put those two facts next to each other:
So the region of the norm table above roughly 145 is thin, and above 160 it is empty. What the table prints up there is not a record of how people scored. It is the smooth curve extended past the last observation, which is a reasonable thing to do and a very different thing from a measurement. Test publishers stop at 160 because that is roughly where extending it further stops being defensible.
Our piece on how rare a high IQ score actually is works the same distribution from the other direction, starting from the rarity and asking what score it corresponds to.
Above 160 the norm table is not a record of how anyone scored. It is the curve drawn past the last person the sample contained.
There is a more concrete limit sitting underneath the statistical one. A Wechsler-type full-scale IQ is assembled from subtests, and each subtest is reported as a scaled score from 1 to 19 with an average of 10 and a standard deviation of 3. Nineteen is therefore exactly three standard deviations above the subtest average — and it is the top of the scale no matter how the raw score was earned.
A test-taker who answers every item on a subtest correctly gets a 19. So does a test-taker who answers enough items to reach the top of that age band's conversion table and no more. The subtest simply has no more difficulty to offer, and the difference between those two people is invisible to the instrument. Stack ten such subtests and the composite inherits the limitation: the test ran out of hard questions before the person ran out of ability.
This is the same reason the Advanced form of a matrix test exists alongside the Standard one — a point we take slowly in what a Raven's score actually is. When everybody at the top gets everything right, a test has stopped discriminating, and the fix is harder items rather than a longer scale.
The obvious fix is to write harder questions and extend the scale upward. Publishers largely do not, and the reason is a trade-off rather than a failure of nerve. Every item added at the top costs testing time, and testing time is the scarcest resource in an individually administered assessment — a full battery already runs well over an hour with a trained examiner in the room. Items that only a few people in ten thousand can answer add almost nothing to the accuracy of the scores everybody else receives, while lengthening the session for every single test-taker.
The norming cost is the larger one. To report accurately at four standard deviations you would need a standardisation sample containing enough people at that level to build a table from rather than to extrapolate one — tens of thousands of participants, individually assessed, stratified to match the population. That is a different order of expense from the samples that exist, in exchange for precision about a group small enough to be identified another way. The honest engineering answer is that an instrument built to place the middle of the distribution accurately cannot also resolve the extreme tail, and the middle is what it is built for.
Publishers do provide ways to report above the standard ceiling, and they are deliberately packaged as supplementary procedures rather than as part of the ordinary scoring. Extended norm tables have been published for Wechsler children's scales to allow reporting above 160 when working with very high-scoring groups, and the Stanford-Binet Fifth Edition includes an extended scoring route intended for the same purpose. Both are documented in the tests' own technical materials, and both come with the caution that the confidence around a score obtained that way is much wider than around an ordinary one.
Almost every very large IQ figure in circulation traces back to a different definition. The original IQ was a ratio: an assessed mental age divided by chronological age, multiplied by 100. A seven-year-old performing like a fourteen-year-old scores 200 on that formula, and the arithmetic is unbounded — it will produce any number you like given a young enough child and a high enough assessed mental age.
Ratio IQ has no fixed standard deviation, so its numbers cannot be compared across ages, across tests, or with a modern deviation score at all. It was abandoned for that reason. Figures attached to historical figures are frequently worse still: retrospective estimates produced by reading biographical material, not scores from any administered test. We traced how those numbers get manufactured and repeated in where celebrity IQ numbers come from.
The reliable tell is simple. If a number is quoted without naming a test, a scale and a date, it is not a test result. It is a claim about a test result, and usually a very old one that has been copied forward.
The first two of those are mechanical, and our tools do them: the percentile calculator turns a score into a rank once you name the standard deviation, and the score converter moves a figure between the SD-15 and SD-16 scales, which is where most of the apparent disagreement between two reported IQs comes from in the first place. If you have never seen a result reported with all three figures attached, our assessment is a short way to find out what that format is supposed to look like.
To see what a score is really worth, convert it rather than admire it: name the standard deviation it was measured on, turn it into a percentile against that scale, and read the range. Our percentile calculator and score converter do both translations, and our own assessment reports a result the way one should be reported.
Find your IQ score now! →The ceiling is not a failure of the tests. It is the tests being honest about what they were built to do, which is to place people accurately across the range where nearly everybody actually falls. Precision at four standard deviations would require a standardisation sample orders of magnitude larger than anyone has ever collected, for the benefit of a handful of people, at the cost of the accuracy that matters for everyone else. A test that reports 194 is not more powerful than one that stops at 160. It is a test that has stopped telling you where its numbers came from.
The major individually administered scales report a full-scale IQ up to roughly 160, which is four standard deviations above the average of 100 on a standard-deviation-15 scale. Some publishers provide extended norm tables or supplementary scoring procedures that report higher for very high-scoring groups, but these are explicitly supplementary and carry a much wider confidence range than ordinary scores.
On a standard-deviation-15 scale, 160 is exactly four standard deviations above average, which on a normal distribution is about the 99.9968th percentile — roughly one person in 31,560. That rarity is the reason it sits at the top of the reportable range: a standardisation sample of about 2,200 people would be expected to contain none at that level.
Those figures nearly always come from ratio IQ, the original formula of assessed mental age divided by chronological age, multiplied by 100. It has no fixed standard deviation, produces unbounded numbers for young children, and cannot be compared with a modern deviation score. Many very high figures attached to historical names are retrospective estimates from biographical material rather than test results at all.
Because subtest scaled scores are set to an average of 10 with a standard deviation of 3, which puts 19 exactly three standard deviations above the subtest average. It is the top of the conversion table regardless of raw performance, so a perfect raw score and a merely excellent one can both return 19. The subtest has no harder items left to separate them.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
On a mean-100, standard-deviation-15 scale, 130 sits at the 97.72nd percentile, which leaves about one person in 44 above it. Here is the rarity at every threshold — and why the figures stop describing real people long before they stop being calculable.
Every properly built test publishes a standard error of measurement alongside the score. On a mean-100, standard-deviation-15 scale it is large enough that a reported 104 and a reported 111 are not reliably different results.

The first IQ scores came from dividing mental age by chronological age. The formula worked reasonably well for children and produced nonsense for adults — which is exactly why every major test replaced it.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →