Average scores climbed for decades, then flattened and in places slipped. None of it shows on a score report, because every test is reset so the average is 100 again — which is why a 1990 score is not a 2026 score.
For most of the twentieth century, each generation scored higher on IQ tests than the one before it. The gains were large, steady and international, and they showed up in countries that had little else in common. That run appears to be over. Since the 1990s, national test data in several wealthy countries has flattened, and in a few places average scores have edged down.
None of this appears on a score report, which is the strange part. IQ tests are not graded like a math exam, where a fixed number of correct answers earns a fixed mark. They are scored against a reference group, and every time a test is rebuilt that reference is reset so the average lands on 100 again. The rise, and now the flattening, is engineered out of the number before anyone reads it.
The pattern is named after James Flynn, the researcher who gathered scattered national datasets in the 1980s and showed the same upward drift almost everywhere he looked. The figure usually quoted is around three points per decade in industrialized countries through the middle of the century. It is a rough average, not a constant: the rate varies widely by country, by era and by the kind of question asked.
What matters more is what that does on the scale. Most IQ scores are built so that 100 is average and one standard deviation equals 15 points. Standard deviation is simply the yardstick statisticians use for how spread out a set of numbers is. Three points a decade is a fifth of that yardstick every ten years, so half a century of gains adds up to roughly a full standard deviation — the distance between an average score and one near the 84th percentile, meaning higher than about 84 of every 100 people in the comparison group.
The gains were also lopsided. General knowledge and vocabulary barely moved. The steepest rises came on the most abstract items: spotting the rule in a sequence of shapes, sorting things into categories, reasoning about a situation that does not exist. Flynn argued that people had not become biologically brighter, but had grown fluent in a style of thinking that modern schooling, work and media demand constantly and that pre-war life rarely required.
The tests did not get easier. The world got better at the kind of thinking the tests were built to measure, and then, in some places, it stopped.
The strongest evidence for a stall comes from an unglamorous source: military conscription records. Norway and Denmark tested nearly all young men for decades, using tests that stayed largely unchanged. That gives researchers something rare — a long, unbroken national series nobody had to assemble after the fact. Those records show gains slowing among men born in the 1970s, then running backwards for men born later.
The picture elsewhere is patchier. An analysis of a large online sample of US adults, tested from the mid-2000s into the late 2010s, reported small declines across most types of question. One category moved the other way: spatial reasoning, the mental rotating and matching of shapes. Online samples are not conscription records, and who chooses to take a test on the internet changes over time. But the direction matched the Nordic data.
The obvious explanation for a national decline is that the nation changed — different migration, different family sizes, different demographics. One Norwegian analysis tested that directly by comparing brothers within the same families, and found the decline there too. Whatever is driving it appears to operate inside households, not only between them.
The honest summary is that the reversal is better documented than it is explained. The leading candidates are all environmental, and all hard to pin down.
What is not on that list is any evidence that the biology of a population changed in thirty years. It could not have, at that speed. Whatever these curves track, it is something about environments, schooling and the tests themselves.
This is where an abstract research argument lands on somebody’s actual paperwork. An IQ test does not measure an amount of anything. It measures a rank. Your answers are converted into a score by comparing them against a norming sample: a large group of people, matched to the population on things like age, sex, education and region, who took the same test under the same conditions at a particular point in time. Norming is that comparison step. The number you receive is a statement about where you sat relative to that group.
So the sample’s date is part of the score, the way a price is meaningless without a year attached. A test normed in 1990 compares you with people as they performed in 1990. If scores in the population drifted upward afterwards, that old test hands out numbers that are too high, because it is grading against a weaker curve. Take the often-quoted three points per decade as an illustration: a test still running on twenty-year-old norms would read something like five or six points high on a 15-point scale. That is not a rounding problem. On a report, it is the difference between average and above average.
It is also why the major clinical tests get rebuilt and re-normed every decade or two, and why professional guidance discourages using an edition whose norms have aged out. And it is why comparing your score with a relative’s from decades ago tells you almost nothing unless you know which test each of you sat, and when.
For most people this is trivia. For some it is not. Eligibility for special education support, for gifted programs, and in some countries for military or civil service roles can hinge on a score landing one side of a threshold. US courts have run into the same problem in capital cases, where a finding of intellectual disability can turn partly on an IQ figure. In Hall v. Florida, the Supreme Court held that a state cannot treat a single score as a rigid cutoff, because every score carries a margin of error.
That last point is the one most consumer tests skip. Every score is an estimate, and a good one arrives with a confidence interval: a range the true value is likely to sit inside, given how precise the instrument really is. On well-built clinical tests that band is commonly a few points either side of the reported figure. A test that hands back a bare 127 — no scale named, no range — is telling you less than it knows.
A score is only as good as the group it was compared against. A current assessment should tell you three things: the scale it uses, where that places you among today’s test-takers, and the margin of error around the figure. If a test returns a single number with no scale and no range attached, there is nothing in it you can check.
Find your IQ score now! →The Flynn effect is one of the more surprising findings in twentieth-century social science, and its apparent ending is one of the more unsettled. The practical lesson has not changed in forty years: an IQ score is a comparison, not a measurement, and a comparison is only as good as the group and the year it was measured against.
Most researchers do not read it that way. The gains were concentrated in abstract, rule-spotting items rather than knowledge or vocabulary, which suggests populations became more practiced at a particular style of reasoning that schooling and modern work demand, rather than biologically different. A century is far too short for population-level biological change of that size.
Treat it as a snapshot rather than a fixed fact about you. It compares you with the norming sample of that test edition, as that group performed at the time. If the norms have since been replaced, the same performance today would usually produce a somewhat lower number. Any comparison across decades needs both the scale and the norming date stated.
It depends on the test and the period, so no single figure applies. As an illustration only: if scores in a population drifted upward at the often-quoted rate of about three points per decade, a test still using twenty-year-old norms would read roughly five or six points high on a 15-point scale. That is enough to shift a score across a reporting band or an eligibility threshold.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
The sturdiest results are narrow and decades old. The 2025 studies that drew the headlines rest on self-reports and small preprints — and none of them measured intelligence at all.
Two people can perform identically on two different online tests and walk away with scores twenty points apart. The scale, the comparison group and the margin of error are what separate an assessment from a quiz.
The correlation between them is one of the sturdier findings in cognitive psychology. How strong it is, and what it implies about training either one, is where the agreement stops.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →