When a large Florida district began testing every second grader instead of waiting for a teacher's nomination, the population it identified as gifted changed sharply — not because more students became capable, but because fewer of them had ever been tested at all.

Every gifted program starts from the same quiet assumption: that the children who get tested are a fair sample of the children who would qualify if tested. A study built around a natural experiment in a large Florida school district shows that assumption fails in a specific, measurable way — and shows how much of a program's demographic makeup is decided before a single test is ever scored.
For decades, most U.S. school districts identified candidates for gifted programs the same way: a teacher, or occasionally a parent, had to refer a child before that child ever sat a cognitive ability test. The test itself might be excellent. The bottleneck was upstream of it. A large Florida district changed that for one cohort, testing every child in a target grade automatically rather than waiting for a referral. Economists David Card and Laura Giuliano compared identification rates before and after the change and published the result — "Universal screening increases the representation of low-income and minority students in gifted education" — in the Proceedings of the National Academy of Sciences in 2016.
The pattern was consistent with what a testing-gap theory predicts and hard to explain any other way. Once every child sat the same test regardless of referral, identification of Black, Hispanic and low-income students rose sharply. Identification of white and higher-income students — who had already been referred at comparatively high rates under the old system — moved far less, because the gap for that group had never been as large to begin with. The overall effect was to narrow the demographic gap in who gets identified, without changing the test, the cutoff, or the definition of "gifted" at all.
A gifted cutoff is not a fixed level of ability in the way a passing grade is. It is a percentile — commonly the 97th or 98th percentile on a cognitive ability test such as the CogAT or the NNAT — meaning a rank against a norm sample of same-age peers, not a raw pass mark. That distinction matters here because a percentile only exists in relation to the people who actually took the test. A child who was never referred cannot be placed on that rank at all, whatever their true standing would have been. Our own percentile calculator works the same way on any score you enter: the rank is only ever computed against the people who sat down and took it.
A child cannot be identified on a percentile they were never tested against. Universal screening does not raise ability — it stops filtering out the children whose ability was never measured in the first place.
Universal screening fixes the referral step. A separate policy tool, used by some states and districts, fixes a different part of the same problem by changing the comparison group instead: local norms rank a student's score against their own school's population rather than against a national reference sample. A school whose average score sits below the national mean will still surface its own highest-scoring students under a local-norm rule, even where those same raw scores would not have cleared a cutoff measured against the national sample. It is a different mechanism from universal screening — it does not touch who gets tested, only who a tested score gets compared against — and districts sometimes combine the two rather than choosing one.
The reason any of this is worth a district's budget is that gifted identification is usually a gateway, not a label on its own. It typically controls access to accelerated coursework, specialized instruction, or additional resources that are otherwise rationed. A testing gap that runs along income or language lines does not just misname a handful of students — it redirects real, scarce educational resources along the same lines, compounding whatever gap in opportunity produced the under-referral in the first place. That is the practical reason a purely administrative-sounding change, like removing a nomination step, shows up in outcomes rather than just in a headcount.
Universal screening is not simply a better policy sitting unused. It costs more per identified student, because it tests every child in a grade rather than a pool a teacher has already narrowed down. Some districts split the difference: screening every child only in specific grades, or on a multi-year cycle rather than annually, to capture most of the referral-gap benefit without testing every student every year. None of that changes the underlying finding — it only changes how much of it a given budget can afford to act on.
None of this is an argument against testing itself — quite the opposite. The difference between an achievement test and a cognitive ability test matters here, because gifted-screening instruments sit closer to the latter: they aim to measure reasoning ability directly, with less dependence on what a child has already been taught, which is part of why they can surface ability in students whose prior schooling was uneven. The test is not the problem this research identifies. Who gets to take it is.
Removing the referral step only helps if the test that replaces it does not reintroduce the same bias in a different form. That is a large part of why universal-screening programs typically choose a nonverbal cognitive ability test rather than a verbal one — an instrument that presents patterns, shapes and sequences rather than vocabulary and reading passages, so that a child who is still acquiring English, or whose early schooling was inconsistent, is not penalized for something the test was never meant to measure in the first place. It is the same design choice behind the CogAT and NNAT batteries discussed elsewhere on this site: reduce how much a score depends on prior exposure to a particular language or curriculum, so the number reflects reasoning ability as closely as a standardized test reasonably can.
Curious what a 97th- or 98th-percentile score actually requires on a real test scale? Take our full assessment to see where a result lands, or use the percentile calculator to convert any score you already hold to its rank against the general population.
Find your IQ score now! →The number that ends up on a district's gifted-program roster is not just a measurement of ability. It is a measurement of ability, filtered through whoever decided a child was worth testing in the first place. Universal screening does not change the first part. It removes the filter.
Testing every student in a grade automatically for gifted-program eligibility, rather than requiring a teacher or parent to refer a child before they can be tested at all.
It mainly changes who gets identified, not how many students in a population are gifted. A 2016 study of a large district's switch to universal screening found it substantially increased identification among previously under-referred groups — low-income, Black and Hispanic students — while identification among already well-referred groups changed far less.
It varies by district and by test, but the 97th or 98th percentile on a cognitive ability test is a common threshold — the same kind of top-of-the-distribution rule used elsewhere in testing, just applied to a school-age norm sample.
Referral depends on a teacher recognizing ability through classroom behavior, which research on identification gaps shows is less reliable for English-language learners, for quiet or highly compliant students, and for children whose families are less familiar with how a referral pathway works.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
The two tests American schools use most to screen for gifted programmes print their results on scales that differ by one point of standard deviation. At the level where cutoffs are set, that moves the same printed number by most of a percentile.
National reading results have dropped again, and around forty per cent of American fourth-graders are now below the NAEP Basic level. That is a fact about an achievement test, and achievement tests are built to show exactly the movement an IQ score is built to hide.

A WISC-V report does not hand back one score. It hands back six — and the headline Full Scale IQ is built from a specific seven of ten subtests, not an average of all five indexes.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →