IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

Gifted Cutoff Scores

Scores & Scales

Gifted Cutoff Scores: Why a Single Number Misses Children

Most gifted programmes draw their line at an IQ of 130. The number is a statistical convention rather than a boundary in nature, and three well-documented features of how the score is produced and used mean a strict cutoff reliably excludes children who belong on the other side of it.

Diagram showing a confidence interval straddling the gifted cutoff of 130, illustrating how two children with different obtained scores can have overlapping true-score ranges

An IQ of 130 is the usual entry requirement for a gifted programme, and it is a convention, not a discovery. It marks two standard deviations above the mean on the scales most tests use, which places it at roughly the 98th percentile. Nothing changes at that point in the distribution. There is no cognitive category boundary there; there is a round number chosen because it is convenient.

That would matter less if the number the cutoff is applied to were exact. It is not, and the gap between an obtained score and the quantity it estimates is large enough to change the answer for a great many children. Three separate features of the process compound: measurement error, the choice of comparison group, and who gets tested in the first place.

Where 130 comes from

Modern tests are scored so that the population mean is 100 and the standard deviation is 15. Two standard deviations up is 130, which leaves about 2.3 per cent of the population above it. That is the entire derivation. The choice of two standard deviations mirrors the convention on the other side of the distribution, where a similar cutoff has historically been used in defining intellectual disability.

Because it is a percentile in disguise, the same label means different things on different instruments. A test standardised with a standard deviation of 16 rather than 15 puts the same percentile at a different number, and older scales computed scores in ways that do not map cleanly onto percentiles at all. Comparing a reported 132 from one instrument with a 129 from another is not a meaningful comparison — the underlying scales differ, as IQ classifications sets out.

The measurement error nobody applies

Every well-constructed test reports a standard error of measurement, and every properly written report expresses the result as an interval rather than a point. On the major individually administered scales, the ninety-five per cent interval around a full-scale score is usually about five points either side.

Work through what that does to a cutoff. A child who obtains 128 has a plausible range running from roughly 123 to 133. A child who obtains 132 has a range from roughly 127 to 137. Those ranges overlap across most of their width. The two children are not meaningfully different on the thing the test estimates, and yet a strict cutoff admits one and refuses the other.

The same arithmetic explains why retesting produces so many reversals. A child who scores 127 in March and 133 in October has not become more able; a second draw from the same distribution came out differently, which is what the interval was warning about. Reading an interval correctly is the single most useful skill for anyone handling one of these reports, and it is covered in how to read an IQ test report.

Diagram showing a confidence interval straddling the gifted cutoff of 130, illustrating how two children with different obtained scores can have overlapping true-score ranges
Diagram showing a confidence interval straddling the gifted cutoff of 130, illustrating how two children with different obtained scores can have overlapping true-score ranges

National norms against local norms

A standard score compares a child to a nationally representative sample. For deciding whether that child needs a different level of instruction than their classmates are getting, the relevant comparison is often the classmates.

The two diverge sharply in schools whose intake is not representative. In a high-achieving school, a large fraction of pupils may clear a national cutoff, so the cutoff stops discriminating and the programme becomes oversubscribed. In a school serving a disadvantaged catchment, almost nobody clears it, so a child who is dramatically ahead of everyone around them and plainly under-challenged is not identified — the label is reporting the catchment rather than the child.

Local norms address this by ranking within the school. They are not a softer standard; they answer a different and more relevant question, namely whether this child is being taught at the right level given the class they are actually in. The two criteria are best used together.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now!

Secure & encryptedInstant results10–20 minutes

The bigger leak: who gets tested at all

Measurement error and norm choice both assume a test happened. In many systems, identification begins with a nomination from a teacher or a parent, and only nominated children are assessed. Every child never nominated is outside the process before any cutoff is applied.

Research on districts that switched from referral-based identification to universal screening — testing every child in a given year group rather than waiting for a nomination — has found substantial increases in the number of children identified from groups that were previously underrepresented, including children from low-income households and those whose first language is not the language of instruction. The children were there; the referral step was not finding them.

This is the largest of the three effects and the easiest to fix, and it is worth being precise about what it shows. It is not a claim that the test was biased against those children. It is a claim that the step before the test decided who would be measured, and that step was not neutral. The instructions given around an assessment can matter as much as the assessment, a theme that also runs through stereotype threat.

The children a cutoff is worst at finding

Two groups are missed so consistently that the pattern is a known feature of the system rather than an accident of any one district.

The first is children whose ability and difficulty coexist — a strong reasoner who is also dyslexic, has attention difficulties, or is autistic. A full-scale score is an average across indexes, and averaging a very high reasoning index with a much lower processing speed or working memory index produces a middling composite that describes neither. The composite is the number the cutoff reads, so the child is refused on the strength of a figure that no clinician would treat as meaningful. Where the index scores are far apart, the full-scale figure should be set aside and the profile read instead — ADHD, autism, dyslexia and IQ test scores covers what those profiles look like.

The second is children still acquiring the language of the test. Verbal subtests measure vocabulary and verbal reasoning in a specific language, and a child two years into learning it will score below their reasoning ability on those subtests and much closer to it on the non-verbal ones. Averaging the two produces a composite that mostly reports language exposure. The gap between the two kinds of subtest is set out in verbal and non-verbal IQ scores.

In both cases the test is doing what it was built to do. The failure is in reducing its output to one number and comparing that number to a line.

What a parent can reasonably do with this

None of this means the score is worthless. It means a single number compared against a single line is a weak decision rule built on top of a reasonable measurement.

  • Ask for the interval, not the number. A proper report gives one. If a decision rests on a point estimate a few points from the line, the interval is the relevant fact.
  • Ask what the comparison group was. National norms and local norms answer different questions, and which one was used should be stated rather than inferred.
  • Ask whether screening is universal. If identification depends on nomination, then not being nominated is not evidence about a child.
  • Treat one session as one session. Attention, sleep and rapport with the examiner all move a result, which is why the same child produces different numbers on different days.

The useful question is not whether a child crosses a line. It is whether they are being taught at a level that fits them, which a score can inform and cannot settle. For what the high end of the scale does and does not mean, see what is genius IQ level; for how age affects the interpretation of a childhood score, see what age is most accurate to take an IQ test.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

How to Read an IQ Test Report

Scores & Scales

How to Read an IQ Test Report: FSIQ, Indexes and Confidence Intervals

An IQ report puts numbers on at least three different scales and almost never explains which of them matters. Here is what FSIQ, the index scores, the subtest scaled scores, the percentile rank and that bracketed range each mean, and when the headline number should not be read at all.

Annotated layout of a typical IQ test report showing the Full Scale IQ with its confidence interval, the index scores on a mean of 100, the subtest scaled scores on a mean of 10, and the percentile rank column

An IQ test report puts numbers on at least three different scales on the same page, and it rarely says so. The Full Scale IQ and the index scores use a mean of 100. The subtest scores use a mean of 10. The percentile rank is on a scale of 1 to 99 and is not a percentage of anything. Read them as though they were comparable and you will draw the wrong conclusion within about thirty seconds.

This guide walks the report from the top down: what each number is, what a normal amount of variation looks like, and the one circumstance in which the headline figure should be set aside entirely. It covers the Wechsler family — the WAIS for adults and the WISC for children — because those account for the large majority of reports people are handed.

Full Scale IQ, and the bracket after it

The Full Scale IQ (FSIQ) is the composite: a single figure derived from the subtests, scaled so that the population average is 100 and the standard deviation is 15. Around two-thirds of people fall between 85 and 115, and about 95 per cent between 70 and 130. Our IQ bell curve page shows the distribution these figures come from.

Immediately after it you will usually see something like 112 (95% CI: 107–117). That bracket is a confidence interval, and it is the most informative and most ignored item on the page. It exists because no test is perfectly reliable: retest the same person and the score moves a little. The interval is the band within which the true score most plausibly sits. For the FSIQ on a modern Wechsler battery it is typically around plus or minus four to five points.

The practical consequence: a 112 and a 116 from the same report are not meaningfully different, and neither are two people three points apart. Treating a single point as informative is the single most common misreading of these reports, and it is the reason our article on what an IQ score really means leads with the range rather than the point estimate.

Index scores: the four or five numbers that matter more

Beneath the FSIQ sit the index scores, also on a mean of 100 and a standard deviation of 15. Each summarises one broad domain. The current Wechsler scales for children and the newest adult edition use five:

  • Verbal Comprehension (VCI) — word knowledge, verbal concepts, acquired verbal reasoning.
  • Visual Spatial (VSI) — construction and analysis of visual material, block-design style tasks.
  • Fluid Reasoning (FRI) — inferring rules from novel patterns, closest to the matrix items on an online test.
  • Working Memory (WMI) — holding and manipulating information over short intervals.
  • Processing Speed (PSI) — how quickly simple, well-defined visual tasks are completed.

Older reports, including those using the WAIS-IV, combine the second and third into a single Perceptual Reasoning Index, giving four rather than five. If your report shows PRI instead of VSI and FRI, that is which edition was administered, not an omission. The domains map closely onto the ability structure described on our page about how IQ tests work.

For most purposes the index profile is more useful than the FSIQ. A single composite averages away exactly the pattern that a referral question usually turns on — strong verbal ability alongside weak working memory says something specific, and it disappears entirely once it is folded into one number.

Annotated layout of a typical IQ test report showing FSIQ, index scores, subtest scaled scores and percentile rank
Annotated layout of a typical IQ test report showing FSIQ, index scores, subtest scaled scores and percentile rank

Descriptors, and strengths that are only relative

Most reports attach a word to each score — Average, High Average, Superior and so on. These are labels for bands on the same scale, they vary between publishers and editions, and newer manuals have deliberately moved away from the older, more stigmatising terminology at the low end. Read the number and the percentile; treat the adjective as shorthand. Our article on IQ classifications sets out how the bands are drawn and why they differ.

One further distinction causes constant confusion. A normative strength means high compared with the general population. A personal or relative strength means high compared with the rest of your own profile. Someone can have a personal strength in verbal comprehension that is still below the population average, and a personal weakness in processing speed that is comfortably above it. Good reports say which sense they mean; not all of them do.

Scaled scores: the numbers between 1 and 19

Individual subtests are reported as scaled scores on a completely different metric: a mean of 10, a standard deviation of 3, and a practical range of 1 to 19. A scaled score of 10 is exactly average. 13 is one standard deviation above; 7 is one below.

The trap is obvious once stated. A subtest score of 12 is a good result, and an index score of 12 would be impossible. If a number in the report is between 1 and 19 it is a subtest and it lives on the mean-of-10 scale — multiply the distance from 10 by five and add 100 for a rough IQ-scale equivalent, so a 13 is roughly the same standing as a 115.

Subtest scores are also the least reliable figures on the page. Individual subtests have narrower reliability than composites, so their confidence intervals are proportionally wider. Modern practice discourages interpreting a single low subtest in isolation, which does not stop people doing it.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now!

Secure & encryptedInstant results10–20 minutes

Percentile rank: the number to quote

The percentile rank says what share of the reference group scored at or below that level. An FSIQ of 100 is the 50th percentile. 115 is roughly the 84th; 130 is roughly the 98th. It is not a percentage score and has nothing to do with how many items were answered correctly.

If you plan to explain a result to anyone, the percentile is the figure to use. It is the one number on the report that means what a non-specialist assumes it means. Our IQ percentile calculator converts in either direction, and the score converter handles the other complication — some tests use a standard deviation of 16 rather than 15, so the same percentile carries a different IQ number.

When the Full Scale IQ should not be reported

There is one genuinely technical rule that reports often apply without explaining. If the index scores differ from each other sufficiently — a large gap between the highest and lowest — the FSIQ is describing an average of things that are not behaving alike, and psychologists will say it should be interpreted with caution or not at all.

In that situation many reports substitute the General Ability Index (GAI), a composite built from the verbal and reasoning indices only, leaving out working memory and processing speed. It is used when those two are depressed by something that is not general ability — attention difficulties, anxiety, motor slowing. Because processing speed is the least g-loaded index, a low PSI can pull an FSIQ down several points while telling you very little about reasoning.

A large split of this kind is common in the profiles discussed in our article on ADHD, autism, dyslexia and IQ scores, which is exactly why the GAI exists.

What else belongs in a real report

A competent report is not just a score table. Look for all of these; their absence is informative:

  • The referral question — what the assessment was actually asked to establish.
  • The test edition and date. Norms age, and an obsolete edition inflates scores.
  • Behavioural observations during testing — effort, attention, fatigue, whether the result is considered a valid estimate.
  • A statement of validity. A report that never says whether the examiner believed the result is not finished.
  • Recommendations tied to the profile rather than to the headline number.

And one thing that should not be there: a diagnosis derived from scores alone. An IQ report describes cognitive performance on a given day. Anything beyond that requires other evidence.

If your report came from an online test

Online tests, ours included, produce a much simpler output and should be read more cautiously. There is no examiner, no behavioural observation and no validity judgement, so the result is an estimate of where you sit rather than a clinical finding. We set out what that does and does not support on our IQ test accuracy page, and what a well-built online test should disclose.

Used for what it is — a normed benchmark rather than an assessment — it is a reasonable starting point. Our free IQ test reports a score against age norms with the percentile alongside it, which is the same pair of numbers you should be reading first on any report.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.