On a mean-100, standard-deviation-15 scale, 130 sits at the 97.72nd percentile, which leaves about one person in 44 above it. Here is the rarity at every threshold — and why the figures stop describing real people long before they stop being calculable.
Every rarity claim attached to an IQ score is a claim about a curve before it is a claim about anybody. Ask how rare an IQ of 130 is and the honest reply begins with a question back: on which scale, and rare compared with whom? Answer both and the arithmetic is exact — about one person in 44 on the scale most modern tests use, not the one in 50 the folklore implies. What the arithmetic will not tell you is how much of that figure is a measurement and how much of it is the shape the scale was given in the first place. The further up the scale you go, the more it is the second.
A modern IQ score is not a quantity of anything. It is a position, printed on a scale that was constructed to have a particular shape. The publisher administers the test to a reference sample — a group assembled to stand in for the population the test is meant for — and then fixes the average of that sample at 100. The spread is set with a standard deviation, which is simply a measure of how far apart people's results typically fall from the average. Most contemporary tests, the Wechsler scales among them, use a standard deviation of 15, and every figure below is on that mean-100, standard-deviation-15 scale unless it says otherwise.
A percentile is the share of that reference sample a score equals or exceeds: sitting at the 90th percentile means nine people in ten scored the same or lower. The conversion from score to percentile runs through the normal distribution, the symmetrical bell-shaped curve the scale was deliberately given.
That last clause carries the whole article. Raw performance does not necessarily arrive normally distributed, and on most modern scales the scoring does not leave the question to chance: scaled scores are typically assigned so that the finished distribution matches the bell curve as closely as the data allow. The curve is an input to the scale at least as much as it is a finding about people. Every rarity figure below is a consequence of that definition, computed exactly, and its exactness says nothing about how well it describes anyone.
One row there is worth saying plainly, because the folklore version of it is repeated so widely: a 130 leaves about one person in 44 above it, not the one in 50 that a rounded-off "top 2 per cent" invites people to infer. For the precise correspondence on any score, our IQ percentile calculator will give it on a named scale rather than from memory — and where the 2 per cent line itself falls, and why Mensa states a percentile rather than a score, we worked through in the piece on Mensa's admission rule. This article stays on the population question, which has a more interesting answer than any single row.
Every figure in that table is a property of the curve the scale was defined by. Not one of them is a count of people.
Read that table down the right-hand column rather than across, and the striking thing is how fast the column accelerates. Fifteen points, from 115 to 130, takes you from about 1 in 6 to about 1 in 44 — a factor of roughly seven. The next fifteen, to 145, multiply the rarity by roughly another seventeen. The fifteen after that, to 160, multiply it by about forty-three. The score scale is linear by construction and the rarity attached to it is emphatically not, which is why a fifteen-point gap near the average and a fifteen-point gap near the ceiling are not comparable quantities at all. The same acceleration is easier to see as a shape than as a table on our IQ bell curve page.
That has a consequence people rarely draw out. Near the middle of the scale, a rarity figure is a reasonably robust summary of where a large number of observed people sat. Near the top, it is a very large number produced by a very small change in position — and a very small change in position is exactly the thing a test cannot promise. Both halves of that sentence get worse the further out you go, and they get worse together.
Norming is the process of building that reference sample and its score table: recruit people to match the target population on age, sex, education, region and whatever else the publisher decides matters, test them under controlled conditions, and derive the conversion from raw performance to scaled score. Standardisation samples for the major individually administered tests are typically on the order of a couple of thousand people, matched to the population on those characteristics. That is expensive, and it is entirely adequate for the middle of the distribution.
It is not adequate for the tails, and the arithmetic that shows why is the arithmetic of the table above. In a sample of 2,200 drawn from a perfectly normal population you would expect about fifty scores above 130 on a standard-deviation-15 scale — plenty to work with. You would expect about three above 145. You would expect 0.07 above 160, which in practice means none, and observing none is what the model predicts whether or not the model is right out there. To observe even ten people above 160 you would need a sample in the region of 316,000. No intelligence test has ever been normed on anything approaching that.
So the published percentile attached to a very high score is not measured. It is extrapolated: the curve is fitted where the data are dense, then continued outward into a region where the sample is silent. That is a defensible thing to do, and it is not the same kind of claim as the one attached to a score of 110 on the same scale. Many test manuals simply stop, refusing to print a figure above roughly 160 on a standard-deviation-15 scale because there is nothing in the sample on which to base one, and extended norm tables exist for some instruments precisely because the ordinary tables run out.
The real distribution is also generally reported not to be exactly normal, and the clearest case is at the bottom. More people are found at very low scores than a smooth bell curve predicts, and the usual explanation offered is that a portion of severe cognitive impairment has specific organic causes — injury, chromosomal conditions, illness — which add cases the curve does not account for. At the top, the corresponding question is unsettled, and it is unsettled for the reason just given rather than because nobody has looked: the samples needed to tell a slightly heavy tail from a slightly light one do not exist.
Rarity is computed from how many standard deviations a score sits above the mean, so it depends on the scale quite as much as on the score. Anyone who has met deviation scoring knows that much. What is less obvious is that the dependence is not proportional: a small change in the standard deviation moves the printed number a little and moves the rarity a great deal.
Converting two scores to the same scale before comparing them is not optional, and our IQ score converter exists for exactly that. A number quoted with no scale attached cannot be assigned a rarity at all — worth remembering whenever one turns up in a profile or a headline. The same discipline dismantles national league tables built from incomparable instruments, a problem we took apart in what country IQ rankings cannot support.
Suppose you hold a properly administered result of 130 on a standard-deviation-15 scale. The rarity attached to it is far less precise than the number looks, and the reason is ordinary measurement error. Test manuals publish a standard error of measurement — an estimate of how much a reported score bounces around a person's true standing from one sitting to the next. On a well-constructed full-scale test it is typically in the region of two to three points on that scale, which puts a 95 per cent confidence band — the range the true score most likely falls in — roughly five points either side of the figure printed.
That much is standard advice, and it is usually left there. Push the band through the rarity table and it stops being a formality. A reported 130 with a band from about 125 to about 135 spans the 95.22nd percentile to the 99.02nd — about 1 in 21 at one end and about 1 in 102 at the other. Ten points of score; a fivefold move in implied rarity. That is not a flaw in the test. It is what happens when a quantity growing this steeply is estimated with any error at all, and it is why a rarity claim is even less quotable than the score it was derived from. Our note on how accurate IQ tests are covers where that error comes from.
If you want a rarity figure you can defend, get it the long way round: sit a properly built test under proper conditions, note the scale it reports on, convert the score to a percentile against that scale rather than against a half-remembered rule, and read the confidence band instead of the point. The result will be less dramatic than the folklore and it will survive being checked.
Find your IQ score now! →The useful conclusion is not that high scores are commoner or rarer than advertised. It is that rarity is not a property a score carries on its own. It is produced by a score, a scale, a reference sample and a model of the population, and the further out you go, the more of the work the model is doing and the less of it the sample is. At 115 on a standard-deviation-15 scale the figure is close to a description of the people who were actually tested. At 160 on that same scale it is very nearly a description of the curve alone. Both are printed with the same confident air, and only one has ever been checked against anybody. If you want to know where a result of your own lands, take a properly scored test and read the percentile rather than the number.
On a mean-100, standard-deviation-15 scale, 130 is exactly two standard deviations above average, which is the 97.72nd percentile. About 2.28 per cent of the reference distribution scores higher, or roughly 1 in 44 people — not the 1 in 50 that a rounded-off "top 2 per cent" implies. The figure is computed from the normal curve the scale is defined by, so it describes the model at least as much as it describes a population.
On a mean-100, standard-deviation-15 scale, about 0.135 per cent of the distribution sits above 145 — the 99.865th percentile, or roughly 1 in 741. That figure is computed from the normal curve the scale is defined by rather than counted in a sample, and it should be treated as a property of the model rather than as a population count.
No. On a mean-100, standard-deviation-15 scale, 160 is four standard deviations above the mean, which is the 99.9968th percentile — roughly 1 in 31,600 by the curve, not 1 in a million. The larger caution is that no standardisation sample is anywhere near big enough to observe an event that rare, so the figure is extrapolated rather than measured.
Because rarity depends on how many standard deviations a score sits above the mean, and different scales use different standard deviations. The number 130 is about 1 in 44 where the standard deviation is 15, about 1 in 33 where it is 16, and about 1 in 9 on the standard-deviation-24 Cattell scale. Convert both scores to the same scale before comparing them.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
On the scale most modern tests use, 120 sits around the 91st percentile. Change the scale and the same number moves. Add the measurement error every test carries and it stops being a point at all.
The tables that circulate as national IQ figures come from one compilation built out of whatever studies existed, in whatever years, on whatever samples. Several of the numbers were estimated rather than measured.
Tables ranking average IQ by degree level circulate constantly and are read as though schooling simply sorts people. Pooled longitudinal evidence puts the correlation at about 0.56, and a meta-analysis of more than 600,000 people found that education also raises scores.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →