Two people can perform identically on two different online tests and walk away with scores twenty points apart. The scale, the comparison group and the margin of error are what separate an assessment from a quiz.
Finding an online IQ test takes seconds. Finding one that will say where its number came from takes considerably longer. Web-based intelligence tests have multiplied far faster than any shared standard for what such a test has to do, and the result is a market in which a fifteen-question quiz and a carefully normed assessment produce the same-looking output: three digits, usually flattering, delivered with no margin of error and no word about who the test-taker was being compared against.
That resemblance is the whole problem. Two people can sit two different online tests, perform about equally well, and walk away with scores twenty points apart — not because either test broke, but because the two used different scales, different comparison groups, or in some cases no real comparison group at all. Nothing on either results page would say so. The number arrives looking like a measurement, and measurements are hard to argue with.
An IQ score is not a count of anything. It is a rank in disguise. Answering thirty-four questions correctly means nothing on its own; the score exists only once that raw performance is compared with the performance of a large group of other people. That group is called the norm sample, and building one is the slow, expensive part of test development. It means giving the test under controlled conditions to a large sample of people chosen to resemble the general population on characteristics such as age, sex, education and region, then working out what a typical result looks like at each age.
Once that reference group exists, raw answers can be converted onto a standard scale. On the scale used by the major clinical tests, the average is set at 100, and the spread is described by something called a standard deviation — in plain words, the typical distance between one person’s score and the average. On those tests one standard deviation is 15 points. If scores follow the familiar bell-shaped curve, roughly two-thirds of people fall between 85 and 115, and about 95 percent between 70 and 130. Those boundaries are not facts about the human mind. They are a decision about how to stretch a ruler, and different tests make that decision differently.
This is where many consumer tests quietly fall apart. The Wechsler tests and modern editions of the Stanford-Binet use a standard deviation of 15. Older Stanford-Binet editions used 16. The Cattell scale, familiar to anyone who has looked into high-IQ societies, uses 24. Identical performance therefore produces very different-looking numbers: a result two standard deviations above average is 130 on a 15-point scale and 148 on a Cattell-style one. Neither is wrong. But a site that reports 148 without naming its scale has told the reader almost nothing, and has handed them a figure that will be misread the moment it is set beside a score from anywhere else.
A number without a scale attached is not a score. It is a decoration.
No cognitive test measures perfectly. Attention wanders, a question set is only a sample of all the questions that could have been asked, and the same person tested twice will not produce the same number. Test developers track this with reliability statistics — measures of how consistently a test produces the same result — and convert them into a standard error of measurement, an estimate of how much a score would bounce around on repeated testing under the same conditions. That is why professionally reported results come as a band rather than a point: the score is given alongside a confidence interval, usually spanning several points either side, meaning the person’s true standing very probably sits somewhere inside that range.
The same discipline applies to percentiles. A percentile is the share of the comparison group a person scored at or above: the 84th percentile means about 84 people in every 100 in that group scored the same or lower. It is a genuinely useful way to make a score intuitive. It is also the number most often abused, because near the middle of the distribution a few raw points move the percentile a long way. Quoting a single rounded percentile as settled fact — you are in the 97th percentile — implies a precision that no test can deliver, online or in a clinic.
Even well-built tests have a further wrinkle to manage. Across much of the twentieth century, average scores on standardized intelligence tests rose over time in many countries — the pattern usually called the Flynn effect, though the rise has not continued everywhere in recent decades. Publishers respond by re-norming: periodically rebuilding the comparison sample so that the average stays anchored at 100 for the current population. A test whose norms were collected long ago, or assembled from whoever happened to visit a website, can drift out of alignment and hand out numbers that are too generous. Inflated scores are rarely the kind of error users complain about, which is one reason they persist.
None of this makes online testing worthless — it makes disclosure the thing to look for. A test worth trusting states its scale, describes its norm sample, and reports the score with a confidence range instead of a single triumphant number. That is the standard IQ Metrics holds itself to on /iq-test/, and readers can check an existing result against it using the /iq-percentile-calculator/ and /iq-score-converter/ tools.
Find your IQ score now! →A test that answers all five is not necessarily excellent, but it is being honest about what it is doing — and honesty is checkable in a way that “scientifically validated” splashed across a homepage is not. A test that answers none of them has not given the reader anything usable. The most reliable signal of quality in this market turns out to be counterintuitive: the more willing a test is to state how uncertain its own result is, the more seriously that result deserves to be taken.
It can give a useful estimate if it is built properly — normed on a large comparison group chosen to resemble the general population, scored on a clearly named scale, timed consistently, and reported with a confidence range. What no online test can do is replace an individually administered assessment supervised by a qualified professional, which is what clinical, educational and legal decisions require.
Usually because they use different scales or different comparison groups. A result two standard deviations above average is 130 on the common 15-point scale but 148 on a Cattell-style 24-point scale, and a test that compares you only with its own past visitors is measuring you against a self-selected crowd rather than the general population. Neither difference is visible on the results page unless the site chooses to disclose it.
It is the range within which a person’s true standing most likely falls, given that no test measures perfectly. Because attention varies and any test is only a sample of possible questions, the same person retested will not produce an identical number. Reporting a band of several points either side of the score is honesty about that fact, not a hedge.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
The sturdiest results are narrow and decades old. The 2025 studies that drew the headlines rest on self-reports and small preprints — and none of them measured intelligence at all.
Average scores climbed for decades, then flattened and in places slipped. None of it shows on a score report, because every test is reset so the average is 100 again — which is why a 1990 score is not a 2026 score.
The exam is shorter, taken on a screen, and the second half of each section changes difficulty based on how you did in the first. Raw right-answer counts stopped mapping onto scores the way they used to.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →