National reading results have dropped again, and around forty per cent of American fourth-graders are now below the NAEP Basic level. That is a fact about an achievement test, and achievement tests are built to show exactly the movement an IQ score is built to hide.
Every time a national score report lands, the same inference follows within a day: scores are down, therefore children are getting less intelligent. The inference is not so much wrong as unlicensed. The results being reported come from a kind of test specifically designed to make change visible, and the number people are silently comparing them against comes from a kind of test specifically designed to hold still.
The National Assessment of Educational Progress, usually called the Nation's Report Card, is administered to nationally representative samples of American students. The 2024 results, released on 29 January 2025, made uncomfortable reading. Average reading scores fell two points at both fourth and eighth grade compared with 2022 – a decline that came on top of the three-point drop recorded between 2019 and 2022, and no state posted a gain in reading at either grade.
A later release in September 2025 covering twelfth grade added science and reading declines at the end of schooling, with twelfth-grade reading now roughly ten points below the first such assessment in 1992. Taken together the picture is of a long slide in reading that predates the pandemic and was made worse by it.
An achievement test asks what a student has learned from a defined body of content. NAEP reports on a scale that is deliberately held constant across administrations, and against fixed achievement levels – Basic, Proficient, Advanced – which describe what students at that level can do. Because both the scale and the standards stay put, a change in the population shows up as a change in the number. That is the entire point of the instrument.
An ability test of the sort that produces an IQ score does something structurally different. It is norm-referenced: your raw performance is converted into a position relative to a reference sample of people your own age, and the scale is set so that the average of that sample is 100 with a standard deviation – a measure of spread – of 15. When the test is re-standardised on a fresh sample, the average is reset to 100 again by construction.
A national decline is what an achievement test is built to reveal, and precisely what an IQ score is built to absorb.
The consequence is worth stating plainly. If the entire population genuinely improved or declined between two standardisations, no individual score report would show it. Everyone would still be measured against their own contemporaries, and the average would still print as 100. This is not a flaw; it is what makes a percentile interpretable. It is also why a score from an old test edition is not comparable with a current one, the problem we set out in why IQ norms expire.
None of this means the two measure unrelated things. Reading comprehension leans heavily on vocabulary, working memory and reasoning, and mathematics achievement leans on quantitative reasoning, so achievement and ability tests share a great deal of variance. In practice the correlation between a broad achievement battery and a full-scale IQ is substantial – large enough that schools have long used one to flag candidates for assessment with the other.
The differences that remain are the ones that matter for interpretation. Achievement is cumulative and instruction-dependent: a child who was not taught something cannot demonstrate it, which is why school closures move achievement scores quickly. Ability tests deliberately minimise dependence on specific taught content, which is why they move more slowly and why they are the instrument used when the question is about a student rather than about a curriculum.
The overlap between the two is large enough that they rank most children similarly, and incomplete enough that a meaningful minority are ranked very differently by them. Those disagreements are exactly the cases educators care about: the child whose reasoning outruns their reading, and the child whose diligent achievement outruns their measured ability. A gap between an achievement result and an ability estimate is a reason to look more closely, not a verdict on either number – and our bell curve page shows how many children sit within a few points of any score, which is usually more than people expect.
If your child brings home a state test result or a NAEP-style summary, read it as a statement about what has been taught and retained, not as an estimate of capacity. If you want the second thing, it requires a different instrument, administered under controlled conditions, and reported with a confidence range rather than as a point. The gap between the two is also why screening for gifted programmes typically uses a cognitive abilities battery rather than the state achievement test, even though the two would rank most children similarly.
The reverse caution applies to the current news cycle. Slower reading development in a cohort is a serious educational problem that deserves the attention it is getting – and it is a problem about teaching, time and text, not a measured decline in reasoning ability. Our reporting on screen time and cognitive test scores in children covers one of the explanations most often offered, and shows how much less settled it is than the headlines suggest.
If you want an ability estimate rather than an achievement summary, take a properly timed, uniformly administered test and read the result with its range. Our percentile calculator converts a score into a position in the population, and the bell curve page shows how many people sit within a few points either side.
Find your IQ score now! →Get those five right and most of the confusion dissolves. Reading achievement in the United States has fallen, measurably and over more than a decade. Whether the underlying reasoning ability of the population has moved is a separate question that a report card cannot answer, and an IQ score – by design – will never show you. If you want to see how a redesigned admissions test changes the meaning of a number, our piece on the digital adaptive SAT works through a related case, and you can take our test if you want an ability estimate of your own.
An achievement test measures what someone has learned from a defined body of content and reports it on a scale held constant across years, so population change shows up directly. An IQ test estimates general reasoning ability relative to a reference sample of the same age, and is re-normed so the average is always 100 – which means the same population change would be invisible on an individual report.
No. NAEP measures achievement in reading, mathematics and science against fixed content standards, and a decline there reflects instruction, curriculum, attendance and much else alongside ability. Because IQ tests are re-normed to an average of 100 at every standardisation, they cannot be used to confirm or deny a population-level change from an individual score report.
Average reading scores fell two points at both fourth and eighth grade against 2022, following a three-point decline between 2019 and 2022. Around forty per cent of fourth-graders scored below NAEP Basic, the largest share since 2002, and about a third of eighth-graders did the same. Fourth-grade mathematics rose two points over the same period.
Substantially, because reading and mathematics both draw on vocabulary, working memory and reasoning. They are not interchangeable, though: achievement depends on what was taught, while ability tests deliberately minimise reliance on specific curriculum content, which is why schools screening for gifted programmes normally use a cognitive abilities battery rather than the state achievement test.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
Average scores climbed for decades, then flattened and in places slipped. None of it shows on a score report, because every test is reset so the average is 100 again — which is why a 1990 score is not a 2026 score.
The exam is shorter, taken on a screen, and the second half of each section changes difficulty based on how you did in the first. Raw right-answer counts stopped mapping onto scores the way they used to.
A Finnish cohort followed 260 children from primary school into adolescence and measured cognition at the end. The children who had spent more time on screens did not do worse. On the processing measures they did slightly better.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →