IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

How to Read an IQ Test Report

Scores & Scales

How to Read an IQ Test Report: FSIQ, Indexes and Confidence Intervals

An IQ report puts numbers on at least three different scales and almost never explains which of them matters. Here is what FSIQ, the index scores, the subtest scaled scores, the percentile rank and that bracketed range each mean, and when the headline number should not be read at all.

Annotated layout of a typical IQ test report showing the Full Scale IQ with its confidence interval, the index scores on a mean of 100, the subtest scaled scores on a mean of 10, and the percentile rank column

An IQ test report puts numbers on at least three different scales on the same page, and it rarely says so. The Full Scale IQ and the index scores use a mean of 100. The subtest scores use a mean of 10. The percentile rank is on a scale of 1 to 99 and is not a percentage of anything. Read them as though they were comparable and you will draw the wrong conclusion within about thirty seconds.

This guide walks the report from the top down: what each number is, what a normal amount of variation looks like, and the one circumstance in which the headline figure should be set aside entirely. It covers the Wechsler family — the WAIS for adults and the WISC for children — because those account for the large majority of reports people are handed.

Full Scale IQ, and the bracket after it

The Full Scale IQ (FSIQ) is the composite: a single figure derived from the subtests, scaled so that the population average is 100 and the standard deviation is 15. Around two-thirds of people fall between 85 and 115, and about 95 per cent between 70 and 130. Our IQ bell curve page shows the distribution these figures come from.

Immediately after it you will usually see something like 112 (95% CI: 107–117). That bracket is a confidence interval, and it is the most informative and most ignored item on the page. It exists because no test is perfectly reliable: retest the same person and the score moves a little. The interval is the band within which the true score most plausibly sits. For the FSIQ on a modern Wechsler battery it is typically around plus or minus four to five points.

The practical consequence: a 112 and a 116 from the same report are not meaningfully different, and neither are two people three points apart. Treating a single point as informative is the single most common misreading of these reports, and it is the reason our article on what an IQ score really means leads with the range rather than the point estimate.

Index scores: the four or five numbers that matter more

Beneath the FSIQ sit the index scores, also on a mean of 100 and a standard deviation of 15. Each summarises one broad domain. The current Wechsler scales for children and the newest adult edition use five:

  • Verbal Comprehension (VCI) — word knowledge, verbal concepts, acquired verbal reasoning.
  • Visual Spatial (VSI) — construction and analysis of visual material, block-design style tasks.
  • Fluid Reasoning (FRI) — inferring rules from novel patterns, closest to the matrix items on an online test.
  • Working Memory (WMI) — holding and manipulating information over short intervals.
  • Processing Speed (PSI) — how quickly simple, well-defined visual tasks are completed.

Older reports, including those using the WAIS-IV, combine the second and third into a single Perceptual Reasoning Index, giving four rather than five. If your report shows PRI instead of VSI and FRI, that is which edition was administered, not an omission. The domains map closely onto the ability structure described on our page about how IQ tests work.

For most purposes the index profile is more useful than the FSIQ. A single composite averages away exactly the pattern that a referral question usually turns on — strong verbal ability alongside weak working memory says something specific, and it disappears entirely once it is folded into one number.

Annotated layout of a typical IQ test report showing FSIQ, index scores, subtest scaled scores and percentile rank
Annotated layout of a typical IQ test report showing FSIQ, index scores, subtest scaled scores and percentile rank

Descriptors, and strengths that are only relative

Most reports attach a word to each score — Average, High Average, Superior and so on. These are labels for bands on the same scale, they vary between publishers and editions, and newer manuals have deliberately moved away from the older, more stigmatising terminology at the low end. Read the number and the percentile; treat the adjective as shorthand. Our article on IQ classifications sets out how the bands are drawn and why they differ.

One further distinction causes constant confusion. A normative strength means high compared with the general population. A personal or relative strength means high compared with the rest of your own profile. Someone can have a personal strength in verbal comprehension that is still below the population average, and a personal weakness in processing speed that is comfortably above it. Good reports say which sense they mean; not all of them do.

Scaled scores: the numbers between 1 and 19

Individual subtests are reported as scaled scores on a completely different metric: a mean of 10, a standard deviation of 3, and a practical range of 1 to 19. A scaled score of 10 is exactly average. 13 is one standard deviation above; 7 is one below.

The trap is obvious once stated. A subtest score of 12 is a good result, and an index score of 12 would be impossible. If a number in the report is between 1 and 19 it is a subtest and it lives on the mean-of-10 scale — multiply the distance from 10 by five and add 100 for a rough IQ-scale equivalent, so a 13 is roughly the same standing as a 115.

Subtest scores are also the least reliable figures on the page. Individual subtests have narrower reliability than composites, so their confidence intervals are proportionally wider. Modern practice discourages interpreting a single low subtest in isolation, which does not stop people doing it.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now!

Secure & encryptedInstant results10–20 minutes

Percentile rank: the number to quote

The percentile rank says what share of the reference group scored at or below that level. An FSIQ of 100 is the 50th percentile. 115 is roughly the 84th; 130 is roughly the 98th. It is not a percentage score and has nothing to do with how many items were answered correctly.

If you plan to explain a result to anyone, the percentile is the figure to use. It is the one number on the report that means what a non-specialist assumes it means. Our IQ percentile calculator converts in either direction, and the score converter handles the other complication — some tests use a standard deviation of 16 rather than 15, so the same percentile carries a different IQ number.

When the Full Scale IQ should not be reported

There is one genuinely technical rule that reports often apply without explaining. If the index scores differ from each other sufficiently — a large gap between the highest and lowest — the FSIQ is describing an average of things that are not behaving alike, and psychologists will say it should be interpreted with caution or not at all.

In that situation many reports substitute the General Ability Index (GAI), a composite built from the verbal and reasoning indices only, leaving out working memory and processing speed. It is used when those two are depressed by something that is not general ability — attention difficulties, anxiety, motor slowing. Because processing speed is the least g-loaded index, a low PSI can pull an FSIQ down several points while telling you very little about reasoning.

A large split of this kind is common in the profiles discussed in our article on ADHD, autism, dyslexia and IQ scores, which is exactly why the GAI exists.

What else belongs in a real report

A competent report is not just a score table. Look for all of these; their absence is informative:

  • The referral question — what the assessment was actually asked to establish.
  • The test edition and date. Norms age, and an obsolete edition inflates scores.
  • Behavioural observations during testing — effort, attention, fatigue, whether the result is considered a valid estimate.
  • A statement of validity. A report that never says whether the examiner believed the result is not finished.
  • Recommendations tied to the profile rather than to the headline number.

And one thing that should not be there: a diagnosis derived from scores alone. An IQ report describes cognitive performance on a given day. Anything beyond that requires other evidence.

If your report came from an online test

Online tests, ours included, produce a much simpler output and should be read more cautiously. There is no examiner, no behavioural observation and no validity judgement, so the result is an estimate of where you sit rather than a clinical finding. We set out what that does and does not support on our IQ test accuracy page, and what a well-built online test should disclose.

Used for what it is — a normed benchmark rather than an assessment — it is a reasonable starting point. Our free IQ test reports a score against age norms with the percentile alongside it, which is the same pair of numbers you should be reading first on any report.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Processing Speed and IQ

Understanding IQ

Processing Speed and IQ: Does Thinking Fast Mean Thinking Well?

Fast thinkers are assumed to be smart thinkers, and there is a real finding underneath the stereotype: people who score higher on IQ tests do respond faster on tasks with no reasoning content at all. But the link is weaker than the folklore, and on a professional battery the speed index is the least g-loaded.

Diagram showing where processing speed sits in the structure of cognitive ability, with simple reaction time, choice reaction time and inspection time on one axis and their differing strength of association with general intelligence shown alongside

Processing speed and IQ are genuinely related, but far less tightly than the stereotype of the quick-witted genius suggests. People with higher IQ scores do respond faster on tasks that contain no reasoning at all — pressing a button when a light comes on — and that finding has held up for over a century. The association is real. It is also modest, and it gets stronger the more complex the speed task becomes, which tells you something important about what is actually going on.

This matters practically, because speed shows up in three different places: as an index on professional batteries, as a time limit on almost every online test, and as the main casualty of cognitive ageing. Those three are often confused with each other.

Reaction time: the oldest finding in the field

The idea that mental speed underlies intelligence goes back to Francis Galton in the 1880s and was revived by Arthur Jensen in the 1970s. The modern version uses a simple apparatus: a light comes on and you press a button. In the simple reaction time version there is one light and one button. In choice reaction time there are several lights and you must press the matching button.

  • Simple reaction time correlates with IQ only weakly — typically around -0.2 (negative because faster means a lower time and a higher score).
  • Choice reaction time correlates more strongly, and the correlation rises as the number of alternatives rises.
  • Variability in reaction time — how inconsistent someone is across trials — often predicts IQ better than their average speed does.

That last point is the interesting one and it is frequently overlooked. Being reliably quick appears to matter more than being occasionally very quick. Reviews of this literature generally place single-task correlations in the low-to-moderate range, with higher values when several speed measures are combined into a composite rather than used one at a time.

Inspection time

A related paradigm removes the motor response entirely. In an inspection time task, two lines of different lengths flash on screen for a very brief interval and you say which was longer. There is no speed of response involved — you can take as long as you like to answer — only how brief a display you can still make sense of. Meta-analyses of this task have reported associations with IQ around -0.5 once corrections for measurement error and restricted range are applied, making it one of the stronger elementary correlates of general ability.

Where speed sits on a professional test

On the Wechsler batteries described in our guide to professional IQ tests, processing speed is one of the four or five index scores that feed the Full Scale IQ. It is measured with tasks like Coding (matching symbols to digits against the clock) and Symbol Search (scanning for a target shape).

Two facts about that index are worth knowing before you read a report:

  • It is usually the least g-loaded of the indices. Verbal comprehension and fluid or perceptual reasoning carry far more of the general factor than speed does.
  • It is the most easily disturbed. Fatigue, anxiety, medication, motor difficulty, vision and simple unfamiliarity with a pencil task all push it down without touching reasoning ability.

That combination is exactly why psychologists sometimes report the General Ability Index alongside the Full Scale IQ — a composite that leaves working memory and processing speed out. If a report in front of you shows one, our guide to reading an IQ test report explains when that substitution is appropriate.

Diagram showing where processing speed sits in the structure of cognitive ability
Diagram showing where processing speed sits in the structure of cognitive ability

Why online tests time you

Almost every online test imposes a limit, and the reason is psychometric rather than dramatic. Without one, reasoning items stop discriminating: given unlimited time, a large majority of test-takers eventually solve a mid-difficulty matrix, so the item no longer separates anyone from anyone. A time limit keeps the difficulty spread usable across the whole range.

The cost is that a timed score blends two things — how well you reason and how fast you work — and different people pay that cost differently. Someone accurate but deliberate loses points that someone quick and careless does not. This is why our timed IQ test is presented separately from the untimed formats in the test hub: they are answering slightly different questions about you.

It is also why speed-accuracy tradeoff advice is not a trick. Working slightly faster than feels comfortable usually helps on a timed test, up to the point where errors start climbing. Our guide to preparing for an IQ test covers where that point tends to sit.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now!

Secure & encryptedInstant results10–20 minutes

Can you train processing speed?

You can improve the score, and that is not the same as improving the ability. Coding-type tasks respond strongly to practice: do one twice and the second attempt is faster, because you have learned the symbol pairings rather than because you now process information more quickly. This is one reason retesting on the same battery too soon inflates a result, as we cover in the practice effect on ability tests.

The best evidence on deliberate speed training comes from large randomised trials in older adults, where speed-of-processing training produced substantial and durable improvement on the trained tasks. What it produced on untrained abilities is far more disputed, and the broader claims made for that work have been challenged repeatedly. The pattern matches the training literature generally — near transfer is well supported, far transfer is not — which is the same conclusion our page on improving your IQ reaches.

A more useful distinction is between things that raise your speed and things that were suppressing it. Sleep debt, illness, anxiety and unfamiliarity with the response format all depress measured speed without touching capacity, which is the mechanism behind test anxiety and IQ scores. Removing them recovers what was already yours; none of them make you faster than your own baseline.

Speed and ageing: the clearest real-world case

Processing speed is the single most age-sensitive part of cognition. It begins declining in early adulthood and continues steadily, and Timothy Salthouse’s influential processing-speed theory argues that much of the apparent age-related decline in reasoning and memory is a downstream consequence of that slowing rather than an independent loss.

Two mechanisms are usually proposed: when operations are slow, early results decay before later ones finish, and fewer operations fit into the available window. Statistically, controlling for speed accounts for a substantial share of the age effect on other abilities — though “accounts for” in a statistical model is not the same as “causes”, and this remains debated.

Because age-normed IQ compares you with people your own age, none of this shows up on a reported score. The number stays flat while the underlying speed changes, which our page on IQ and age charts directly.

Does thinking fast mean thinking well?

Not reliably, no. Speed is one contributor to a score among several, it is the weakest of the index-level contributors, and the deliberate-and-accurate profile is common among high scorers. Worth separating clearly:

  • Speed of elementary processing — reaction and inspection time — is modestly related to general ability and is not something you can train up meaningfully.
  • Speed on a clerical task — the WAIS speed index — is only loosely related to reasoning and is easily depressed by things that have nothing to do with intelligence.
  • Speed of judgement — answering fast in conversation or in a meeting — is barely studied, closer to a personality trait, and is often the opposite of good thinking. That is the theme of why smart people make bad decisions.

If you want to see how the two behave for you specifically, take our free IQ test once under its normal limit and note where you were rushed. The gap between what you answered and what you could have answered is your own speed contribution, and it is usually smaller than people fear.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Does IQ Predict School Success?

Research & Evidence

Does IQ Predict School Success? What Fifty Years of Data Show

IQ predicts school achievement better than any other single measure psychologists have found, and still leaves roughly three-quarters of the variation unexplained. Both halves of that sentence get ignored by somebody. Here are the actual correlations, what they mean for one child, and what accounts for the rest.

Chart showing that a correlation of 0.5 between IQ and school achievement accounts for about 25 per cent of the variation, with the remaining 75 per cent attributed to prior knowledge, conscientiousness, motivation, teaching quality and circumstance

IQ predicts school success better than any other single measure psychologists have found, and it still leaves most of the variation unexplained. Correlations between cognitive ability scores and school achievement usually land around 0.5, sometimes as high as 0.7 for standardised achievement tests in younger samples. That is a strong result by the standards of social science. It also means roughly three-quarters of the differences between children are about something else.

Almost every argument about testing in schools comes from taking one half of that sentence and dropping the other. Advocates quote the strength of the correlation to justify selecting children by it; critics quote the unexplained remainder to argue the measure is worthless. Both are reading a real number as though it answered a question it does not address.

What a correlation of 0.5 actually means

Square it. A correlation of 0.5 corresponds to about 25 per cent of the variance in the outcome being statistically associated with the predictor. Three-quarters is not.

Concretely: among children with the same IQ score, school achievement still varies enormously. The relationship is strong enough to be visible in a group of five hundred and far too weak to be reliable about any one of them. This is the single most common misreading of the literature — treating a solid group-level correlation as an individual-level forecast.

Chart showing that a correlation of 0.5 between IQ and school achievement accounts for about 25 per cent of the variation, with the remaining 75 per cent attributed to prior knowledge, conscientiousness, motivation, teaching quality and circumstance
Chart showing that a correlation of 0.5 between IQ and school achievement accounts for about 25 per cent of the variation, with the remaining 75 per cent attributed to prior knowledge, conscientiousness, motivation, teaching quality and circumstance

The numbers, outcome by outcome

The correlation is not one figure. It depends heavily on what is being predicted.

  • Standardised achievement tests: about 0.5 to 0.7. The strongest relationship, and unsurprising — achievement tests and ability tests share format, timing and reasoning demands.
  • School grades: about 0.4 to 0.5. Consistently lower, for reasons worth a section of their own.
  • Years of education completed: around 0.5. Reasonably strong, but heavily entangled with family circumstances.
  • Performance within a selective university: weak. Once a group has been filtered on ability, the remaining range is narrow and the correlation shrinks accordingly.

Age matters too, and in a direction that surprises people. The correlation between measured ability and achievement tends to be strongest in the primary years and to weaken through secondary school and beyond. Part of that is restriction of range as cohorts get filtered; part is that accumulated subject knowledge, study habit and choice of subject increasingly dominate as the material gets more specialised. A reasoning test predicts best when there is least to have already learned.

That last point is restriction of range, and it explains a great deal of apparently contradictory research. A predictor always looks weaker inside an already-selected group. The same effect appears when cognitive tests are used in hiring, where the candidate pool has usually been screened already.

Why grades track IQ less closely than test scores

A grade is not a measurement of what a student knows. It is a composite of what they know, whether they handed it in, whether they attended, how they behaved, and a teacher’s judgment of all of that.

The best-known study on this point followed eighth-graders and found that a measure of self-discipline outpredicted IQ for report card grades by a substantial margin — while IQ remained the better predictor of standardised achievement test scores in the same children. Both findings are in the same paper, and quoting either one alone misrepresents it. Conscientiousness wins where sustained daily compliance is measured; ability wins where a single unfamiliar reasoning task is measured.

What accounts for the other three-quarters

  • Prior knowledge. The strongest predictor of learning something new is usually how much of the surrounding subject you already know.
  • Conscientiousness and self-discipline. Homework completion, attendance, deadline behaviour.
  • Teaching and school quality. Large effects, unevenly distributed.
  • Family circumstances. Books, quiet space, illness, stability, adult time.
  • Motivation and interest. Also partly a consequence of earlier success, which makes the causal arrows circular.
  • Test conditions on the day. Sleep, anxiety and illness all move scores, as the evidence on test anxiety shows.

A word on grit specifically, since it is often offered as the answer. Meta-analytic work suggests grit is largely a relabelling of conscientiousness and adds modest incremental prediction of academic performance beyond it. The broader point survives; the specific construct is weaker than its popularity implies.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now!

Secure & encryptedInstant results10–20 minutes

The confound that will not go away

Socioeconomic status correlates with measured ability and with school achievement independently, which makes every simple comparison ambiguous. A family with more resources supplies more books, quieter study space, better nutrition, more adult conversation, less disruption from moving house or illness, and more experience of the conventions a test is written in. All of that raises the score and the achievement at once.

Studies control for it statistically, and controlling is not the same as removing it. The variables available in a dataset — parental income, parental education, an area deprivation index — are crude stand-ins for a diffuse advantage, so residual confounding is the norm rather than the exception. Treat any estimate of the ability-achievement link as an upper bound on the causal contribution of ability, not a measurement of it.

There is one finding worth flagging because it is frequently over-read in both directions: several studies report that the heritability of cognitive ability appears lower in more deprived environments, which would imply that circumstance constrains what ability can express. The result has replicated in some samples and failed to in others, and it should be held loosely. The wider evidence on inheritance is set out in what twin studies actually show.

Causation runs both ways

It is tempting to read all this as ability causing achievement. The arrow is not one-directional.

Natural experiments exploiting changes in compulsory schooling laws find that additional years of education raise IQ scores, with estimates commonly in the range of one to five points per year. Schooling does not merely reveal ability; it partly builds the thing the test measures. That is consistent with the twentieth century’s rising scores described in why IQ norms expire, and with the careful account of what can and cannot be changed in whether you can improve your IQ. It also sits alongside the genetic evidence rather than against it: heritability describes variation in a population under given conditions, and says nothing about how much a score would move if the conditions changed.

The practical consequence is that a low score in a child who has missed a great deal of school is not a stable trait measurement. It is a reading taken partly on the schooling. Repeating the assessment after a period of consistent attendance is not redundant — it is the only way to separate the two.

What this means for one child

Four things follow, and they are the practical payoff of everything above.

  • A single score is a snapshot with an error band. Childhood scores are less stable than adult ones, as how scores change over time sets out.
  • Extreme scores drift toward the average on retesting, for statistical reasons rather than psychological ones — regression to the mean is the mechanism.
  • A test taken in a second language, or under an unfamiliar format, measures partly those things — the practical upshot of cultural bias in IQ tests.
  • Predicting a group is not forecasting a person. A 25 per cent variance share is a headwind or a tailwind, not a destination.

What follows for schools

The prediction is good enough to be useful for allocating support and poor enough to be dangerous for allocating opportunity. Those two uses look similar and are not. Using a score to decide which children get extra reading help is a low-cost decision that is easily reversed if wrong. Using the same score to decide which children enter an academic track at eleven is a high-cost decision that is difficult to reverse and that compounds over years.

The historical case against selection by test at a fixed age was built on exactly this asymmetry rather than on the tests being meaningless, and how testing has been used in education and employment traces where that argument went.

The defensible summary

Cognitive ability is a real and useful predictor of school achievement, the best single one available, and a poor basis for deciding what any individual child will do. Those statements are consistent, and holding all three at once is the whole skill.

If you want a score read the way this article argues it should be — with its scale, its percentile and its confidence range stated rather than a bare number — that is what our IQ test reports, and the version for children is normed against age-matched peers rather than adults.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Verbal vs Non-Verbal IQ Scores

Scores & Scales

When One Half of Your IQ Score Is 15 Points Higher Than the Other

Two people can hold the same IQ number and have completely different profiles, and a split between the verbal and non-verbal indexes is far commoner than most test-takers expect. Here is what base-rate tables and compounding measurement error say about how large a gap has to be before it means anything.

Chart showing a verbal index of 118 and a non-verbal index of 103, each with its measurement error band, and the 15-point gap between them carrying a much wider band of roughly 4 to 26 points

A gap between your verbal and your non-verbal scores is far more common than most people assume. Published base-rate tables from the major batteries show index splits of 10 to 15 points turning up in a substantial minority of the standardisation sample, so a verbal vs non-verbal IQ score difference of that size is usually unremarkable. The useful question is not whether you have a split, but whether yours is big enough to survive the measurement error sitting on both sides of the subtraction.

One number is an average, and averages destroy shape

A full-scale score is a weighted average of several index scores, and averaging is a lossy operation. Two people can arrive at the same 112 by completely different routes: one of them even across every domain, the other with a strong verbal index dragging a weaker visual-spatial one up to the mean. The composite is identical; the profiles are not.

That is why psychologists read the indexes before the total, and sometimes decline to interpret the total at all when the indexes disagree sharply. If the parts of a measure point in different directions, their average summarises something that does not really exist as a single quantity.

The two directions a split can run

Index score discrepancies are not symmetrical in what they suggest. A verbal score sitting above a non-verbal one keeps different company from the reverse pattern, and neither is a diagnosis. Treat the two shapes described here as typical associations, nothing stronger.

Chart showing a verbal index of 118 and a non-verbal index of 103, each with its measurement error band, and the 15-point gap between them carrying a much wider band of roughly 4 to 26 points
Chart showing a verbal index of 118 and a non-verbal index of 103, each with its measurement error band, and the 15-point gap between them carrying a much wider band of roughly 4 to 26 points

Verbal sitting above non-verbal

This shape is common in people with long formal education, heavy reading habits and verbally loaded work: teaching, law, journalism, anything whose day is made of words. Vocabulary and general knowledge keep accumulating with exposure, while timed matrix and pattern tasks do not reward exposure in the same way, so the pattern tends to flatter older test-takers.

That asymmetry between accumulated knowledge and reasoning on the spot has a standard name in the literature, and how IQ tests are built and scored sets it out. The point here is only that a verbal edge is often a record of practice rather than a difference in raw horsepower.

Non-verbal sitting above verbal

The reverse split is common among people testing in a second language, people whose schooling was interrupted or happened in a different system, and younger adults, who have simply had fewer years to accumulate the vocabulary and general knowledge the verbal side rewards. Reading and language difficulties can also hold down verbally loaded scores without that implying a matching limit on reasoning, which is a clinical question rather than a scoring one and is handled in the piece on ADHD, autism, dyslexia and test scores.

Language load is the obvious confound in this direction, and batteries differ in how much of it they carry; those differences are compared in the guide to professionally administered IQ tests.

How big does a gap have to be before it is signal?

Two things have to hold before a split is worth interpreting. It has to be unusual relative to the population, and it has to be larger than the combined error of the two measurements that produced it. Most of the gaps people worry about fail both tests.

Base rates: lopsided profiles are ordinary

Test manuals publish base-rate tables showing how often each size of index difference occurred in the standardisation sample. The consistent finding is that double-digit splits are common among entirely typical people: published tables put differences of roughly 10 to 15 points in a substantial minority of the sample. The exact percentage depends on the battery and on which pair of indexes you are comparing, so treat any single figure you see quoted as specific to one test rather than universal.

The practical consequence is deflating. A 12-point split is not a finding. Differences of that order turn up often enough in the standardisation samples that seeing one tells you almost nothing about the person holding the score. Gaps only start to look genuinely uncommon well beyond the double-digit range, and where that line sits is decided by the table for that specific index pair on that specific battery, not by any single number that carries across tests.

Why subtracting two noisy numbers widens the error

Every score carries measurement error. A well-built battery with a reliability near 0.90 has a standard error of measurement of roughly 4 to 5 points, which is why a score of 115 is honestly reported as a roughly 95 per cent band of about 106 to 124 rather than as a single point; that arithmetic is worked through in the piece on how accurate IQ tests really are. Index scores rest on fewer items than the full scale, so their bands are if anything a little wider than that.

Now subtract one of those bands from another. Uncertainties do not cancel when you take a difference, they accumulate: the gap between two index scores is a noisier quantity than either index on its own, so its error band is wider than the band on either score you started with. Two errors of 4 to 5 points combine into an error on the difference of roughly 6 to 7 points, which puts a band of something like 12 or 13 points either side of any gap you measure. A difference has to clear a higher bar than a score does before it can be told apart from zero.

Said plainly: if each of two scores can land several points either side of your true standing, a gap of 8 or 10 points between them can be manufactured by nothing more than which day you sat down. Small splits are usually not small effects but no effect at all, which is why a modest gap so often shrinks or flips direction on a retest.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now!

Secure & encryptedInstant results10–20 minutes

Subtest scatter is even weaker evidence

Once a split appears, the temptation is to go one level deeper and interpret individual subtests: strong on one, weak on another, therefore a profile. Resist that. Subtests are shorter than indexes, shorter means less reliable, and less reliable means the scatter across them is noisier still.

Decades of research on interpreting individual subtest peaks and troughs have been unkind to the practice. Index-level differences are the smallest unit worth taking seriously, and only when they clear both the base-rate hurdle and the error hurdle. Keep constructs separate too: a weak memory-loaded score is not the same evidence as a weak reasoning score.

What a real discrepancy actually licenses you to conclude

Suppose your split profile clears every hurdle: large, rare in the base-rate table, and still there on a second sitting. What have you learned? A hypothesis about how you learned and where you are practised — not a second IQ, and not a diagnosis.

  • It is descriptive, not causal. A high verbal index says the vocabulary and verbal reasoning are there. It does not say whether they came from schooling, reading, work or something else entirely.
  • It is not a ceiling on the weaker side. Index scores move with familiarity and practice, and retest gains are usually larger on the non-verbal side than on the verbal one.
  • It is not a label. Score patterns are consistent with many explanations and specific to none; a diagnosis needs history, observation and criteria that no test result supplies.
  • It is not a career instruction. The link between profile shape and occupational outcomes is far too loose to guide an individual decision.

The honest use of a split is narrower. It tells you which kind of task you are likely to find comparatively effortful, which is worth knowing when you decide how much preparation an exam or an aptitude screen deserves.

How to read your own rough profile

Breaking the composite apart is the only way to see shape, and a few rules keep the exercise honest.

  • Sit the domains separately: the verbal reasoning test, the spatial reasoning test and the numerical reasoning test, ideally on different days, so that fatigue does not manufacture a gap for you.
  • Convert every result to a percentile first. Different tests are not marked on the same ruler, and percentiles are the closest thing to a common currency.
  • Compare the percentiles, not the raw scores. If your two most extreme domains land within about ten percentile points of each other, treat the profile as flat — flat is the normal answer, and anything wider still has to clear both hurdles above before it means anything.
  • Re-sit your two most extreme domains once. A gap that survives a repeat is worth thinking about; a gap that moves was noise wearing a costume.

One caution about self-administered results: unsupervised tests vary in quality and in how carefully they were normed, so a split measured across two different tests carries all the error above plus the gap between two norming samples. Comparing within one family of tests is safer.

Where to start if your number feels uneven

If you are holding one score and a suspicion that it does not describe you evenly, the cheapest next step is to take it apart rather than to take it again. Sit the domains one at a time, run each result through the IQ percentile calculator, and look at the shape instead of the total.

Most people find their profile is flatter than they expected, and that is the good outcome: a flat profile means the composite you already have is doing its job. A large, repeatable split is rarer, and even then it describes where your practice has gone rather than what you can learn next.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

ADHD, Autism and Dyslexia in IQ Testing

Taking a Test

How ADHD, Autism and Dyslexia Change What an IQ Test Measures

A test score is a measurement, and a measurement can be biased by the conditions under which it is taken. ADHD, autism and dyslexia each depress performance on particular kinds of item without lowering the reasoning underneath. Here is what actually interferes, condition by condition, and what to do next.

Grid rating how far dyslexia, ADHD and autism affect four IQ test domains: verbal comprehension, processing speed, working memory and matrix reasoning, which is largely spared

If you have ADHD, autism or dyslexia, a score from a timed online test is not meaningless, but for many people it reads lower than their reasoning does. ADHD and IQ test scores interact through the format rather than through intelligence: attention, reading load and the clock all sit between your thinking and the number at the end. What follows names the mechanism for each condition, and what to do about it.

This site publishes tests and explains scores; it has no clinical standing, and nothing here is diagnosis or medical advice. No IQ score, and least of all an unproctored online one, can indicate or rule out ADHD, autism or dyslexia. A qualified assessor is the only route to that answer.

A Depressed Score Is Not a Lower Ability

Much of what is written about neurodevelopmental conditions and testing quietly folds two different claims into one. The first is that measured composites in these groups often come out lower than in comparison samples. For ADHD and dyslexia that is a reasonably consistent finding; for autism it depends heavily on which test was used and how the sample was recruited.

The second claim is that these groups reason less well. That does not follow from the first. A test score is a measurement, and any measurement can be distorted by the conditions under which it is taken. It is more precise to say the score is depressed than that the ability is lower.

A bathroom scale on a thick carpet reads wrong; the reading changed, the weight did not. Much of what these three conditions do to a test result is carpet.

Ordinary test-day variables work the same way; the same person can land several points apart across a fortnight. Our longer piece on the factors that affect IQ test results covers sleep, anxiety, caffeine and practice. The difference is that these three conditions are stable features of how you process information, not the state you turned up in.

Where the Interference Actually Lands

A modern battery is not one thing. It samples several fairly distinct abilities and averages them into a single composite, and interference rarely hits all of them equally. It lands hard on one or two, and the composite carries the damage without saying so.

  • Verbal comprehension – vocabulary, similarities, general knowledge; all of it delivered in written language.
  • Fluid reasoning – matrices and number series, the part closest to raw pattern-finding.
  • Working memory – holding material in mind and manipulating it while you use it.
  • Processing speed – how quickly you get through easy items, not how hard the items are.
Grid rating how far dyslexia, ADHD and autism affect four IQ test domains: verbal comprehension, processing speed, working memory and matrix reasoning, which is largely spared
Grid rating how far dyslexia, ADHD and autism affect four IQ test domains: verbal comprehension, processing speed, working memory and matrix reasoning, which is largely spared

This is exactly why a profile beats a composite, and why the index-by-index breakdown produced by professionally administered tests is worth far more to a reader in this position than the headline figure.

Dyslexia: The Reading Load Hidden Inside the Test

Dyslexia is a difficulty with accurate or fluent word recognition, decoding and spelling, neurobiological in origin and not explained by a lack of instruction. On a test it lands in two predictable places. Verbal comprehension items have to be read, and so do the instructions and the answer options.

Processing-speed items are trivially easy in content and hard to finish. On a professional battery they use symbols rather than words, and are still often lower in dyslexic profiles because naming things quickly is part of what they measure. On an online test the speeded items are usually words and sentences, which puts the reading load back on top of the clock.

Fluid reasoning typically holds. A dyslexic reader who grinds through a vocabulary section will often complete matrix items at roughly the level the rest of their profile predicts, because a matrix does not have to be read. The composite then averages one fair estimate with one depressed estimate and understates the reasoning.

How large that gap gets varies from person to person, and a gap is not automatically meaningful on its own. The companion piece on verbal and non-verbal score splits covers when a difference between two indexes is big enough to interpret at all.

ADHD and IQ Test Scores: Inconsistency, Not a Lower Ceiling

The characteristic pattern for ADHD on a test is not uniformly weaker performance. It is variability. The same person can solve a hard matrix item in seconds and then miss three easy ones in the block that follows, and where in the session an item fell can matter more than how hard it was.

Three separate things cost points here, and they are worth keeping apart because they respond to different fixes.

  • Time pressure. A per-item or whole-test clock converts a slow start into lost items, even when the reasoning would have arrived a few seconds later.
  • Sustained-attention drift. Dozens of near-identical matrix puzzles in a row is close to a worst case: low novelty, no feedback, and wrong answers that feel exactly like right ones.
  • Working-memory load. Items that require holding several constraints at once are harder to finish when the holding itself takes effort.

That last one is worth holding apart from reasoning itself: the two are related but not interchangeable, and someone can have a genuinely low working-memory index and strong fluid reasoning at the same time. An item that loads both is limited by the weaker of the two.

Sleep and time of day move ADHD test performance, and so can where someone is in their usual routine on the day. That is another reason to treat one sitting as a single sample rather than a verdict, and to be sceptical of any score taken at the end of a long day.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now!

Secure & encryptedInstant results10–20 minutes

Autism: Uneven Profiles and Confounded Instructions

There is no autistic IQ profile. Autistic people cover the entire measured range, and the variation inside the group is far larger than any average difference between it and everyone else. Anything that describes one shape of profile as typical is describing a subset and calling it a population.

One pattern does show up often enough to be worth naming: a split between verbal and perceptual performance, running in either direction. Some autistic people score markedly higher on matrix reasoning than on verbal subtests, particularly where verbal items reward conventional usage; others show the reverse.

The larger measurement problem is often the wrapper rather than the item. Instructions written loosely can be read literally and answered correctly by their own logic. Asking which one does not belong has several defensible answers, and an item lost that way costs the same as an item nobody could solve. The score cannot tell the two apart.

Conditions matter too. An unfamiliar room, harsh lighting, the social weight of being watched while you think: none of that is reasoning, and all of it can move a number.

What an Online Screener Does to These Profiles

An unproctored online test has no way to accommodate anyone. It cannot verify a diagnosis, grant extra time, offer a quiet room, read an item aloud, or notice that you stopped concentrating twenty minutes ago. It records what happened and scores it.

Format still changes the size of the problem. A timed IQ test stacks the effects above into their worst configuration: a speed penalty, no way back from a slow opening, and a hard stop that arrives whether or not you were nearly there.

If the clock is your main obstacle, the untimed classical format is a fairer instrument for you. If reading load is the obstacle, a culture-fair test built from figures rather than sentences takes most of the language out from between you and the item. Neither is a formal accommodation; each just removes one confound.

Accommodations, and Who Can Authorise Them

Formal assessment does have answers to all of this, and they are unglamorous rather than clever.

  • Extended time, commonly one and a half times the standard limit, sometimes double.
  • Scheduled breaks between subtests, or the session split across more than one day.
  • A separate quiet room, with the examiner present and nobody else.
  • Untimed administration of subtests where speed is not the ability being measured, reported alongside the standard scoring rather than in place of it.
  • Items read aloud, or answers given orally rather than written.

None of that is self-service, which is the practical point most articles skip. Accommodations are authorised by whoever owns the assessment: a qualified psychologist for a clinical evaluation, a school’s special educational needs process for a child being tested, a university disability service, or an employer’s occupational-health route at work. Each normally wants documentation first, which is one more reason a screener result cannot start the process.

If You Think Your Score Is an Underestimate

Start by treating the number as information about a sitting rather than about you. Which sections felt bad, and why? Running out of clock, losing the thread halfway through a block, and reading the same sentence four times are three different failures, each pointing at a different index dragging the composite down.

Then take the test again in a format that removes the thing you identified, and compare. Two scores with a known difference between the conditions that produced them tell you more than one score of any size, and that comparison is the closest thing to self-accommodation a free screener can offer.

If a defensible number actually matters — for a school placement, an accommodation request, an assessment — an online result will not do that job. What a free online IQ test gives you is a rough bearing and, taken twice under deliberately different conditions, a sense of how much of your score is about the format rather than your reasoning.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.