IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

Spearman's Law of Diminishing Returns: Does g Weaken at High Ability?

Cognitive tests correlate less strongly with each other among high scorers than among low scorers. If that holds, a single full-scale figure carries less information the higher it goes — and the evidence for it is real but smaller and shakier than the idea's popularity suggests.

Spearman's Law of Diminishing Returns: Does g Weaken at High Ability?
Illustration generated for IQ Metrics. No photograph is used. IQ Metrics

What you need to know

  • Charles Spearman observed in 1927 that the correlations between different mental tests are weaker among more able people than among less able people. The pattern is now called Spearman's Law of Diminishing Returns, usually shortened to SLODR.
  • If it holds, one general factor accounts for less of the difference between two high scorers than between two low scorers — so a single full-scale IQ figure summarises a high performer less completely, and the shape of their subtest profile carries more of the information.
  • The evidence is real but contested. A widely cited 1989 analysis by Detterman and Daniel reported subtest intercorrelations in the region of twice as large in the lowest-ability group as in the highest. Later work using better statistical controls has generally found the effect in the same direction but considerably smaller, and some analyses do not recover it at all.
  • The practical consequence is modest and specific: above roughly the 98th percentile, treat a full-scale score as a summary that hides more than it does at the middle of the distribution, and read the index scores underneath it.

Two people take the same battery. One scores 85 on a mean-100, standard-deviation-15 scale; the other scores 135. Both walk away with a single number that is supposed to describe how they did across every part of the test. There is a century-old argument in psychometrics that those two numbers are not doing equally good jobs — that the higher one is a worse summary of the person it describes. The claim is called Spearman's Law of Diminishing Returns, and it is one of the few ideas in intelligence research where the popular version and the evidence have drifted noticeably apart.

The observation Spearman made

Charles Spearman is the reason anyone talks about a general factor at all. His central finding, first set out in 1904 and developed through the 1920s, was that people who do well on one kind of mental test tend to do well on the others — vocabulary, arithmetic, spatial puzzles, memory span. Every pair of tests correlates positively, a pattern known as the positive manifold, and the shared variance is what the letter g refers to. Our explainer on what the g factor actually is covers that ground in full.

In his 1927 book The Abilities of Man, Spearman added a qualification that is far less often quoted. The strength of those correlations, he noted, was not constant across the range of ability. Among people who did poorly, the tests hung together tightly. Among people who did well, they came apart. He described the general factor as behaving like a resource that mattered enormously when it was scarce and mattered progressively less once there was plenty of it — hence the name later attached to the idea.

What "diminishing returns" means in practice

A correlation between two subtests describes how well one predicts the other. If verbal comprehension and matrix reasoning correlate at 0.55 in a group, knowing one score tells you a fair amount about the other. If they correlate at 0.30, it tells you much less. Spearman's claim is that the first figure describes low scorers and the second describes high scorers — and if that is right, several things follow:

  • The general factor explains a smaller share of the differences between able people than between less able people.
  • A full-scale score, which is essentially a weighted average of the parts, therefore loses more information at the top of the range than in the middle.
  • Two people with the same high full-scale figure may have arrived there by genuinely different routes — one on verbal strength, another on spatial — more often than two people with the same average figure.
  • Selecting on a single cut score does less to make a group homogeneous the higher the cut is set.

At the middle of the distribution one number is a fair summary of a person. At the top it is increasingly a summary of an average of things that have started to diverge.

The evidence, which is genuinely mixed

The modern literature starts with a 1989 study by Douglas Detterman and Mark Daniel, who split standardisation samples by ability and compared the subtest correlation matrices. In the lowest-ability groups the average intercorrelation was in the region of twice what it was in the highest-ability groups. That result is the one usually cited, and it is a large effect — large enough that if it were the whole story, test manuals would need to say so on the first page.

It is not the whole story. Splitting a sample by ability and then measuring correlations inside each slice introduces a well-known statistical artefact: restricting the range of a variable reduces its correlations mechanically, whether or not anything about the underlying structure has changed. Much of the work since has been an argument about how to separate a real change in structure from that artefact. Approaches that model the change continuously rather than by cutting the sample into groups — moderated factor analysis, and methods that estimate the factor structure locally along the ability range — generally do find the effect in Spearman's direction, but smaller than the 1989 figures imply. A 2017 meta-analysis by Blum and Holling in Intelligence concluded that support for the law exists across a substantial body of studies while being moderated by sample and method. Other analyses, including work on how abilities differentiate across the lifespan, have found age effects considerably clearer than ability effects.

This is a hard thing to measure, and that is the honest summary of the field rather than a hedge. The quantity in question is not a score but the shape of a correlation matrix, estimated separately in subgroups that are by construction less variable than the whole sample. Tests also run out of difficulty at the top — a battery normed for the general population has few items hard enough to separate people at the 99th percentile, and a ceiling compresses scores in exactly the region where the law predicts something interesting. Some of the reported effect is structural; some of it is the instrument running out of room.

Why it matters if you have a high score

The practical reading is narrower than the idea's reputation. Nobody serious claims g stops existing at high ability, or that a 135 is uninformative. The claim is about how much of the story one number tells, and it bears on a specific set of decisions:

  • Above roughly the 98th percentile, the index scores underneath the full-scale figure deserve more weight than they do in the middle of the range — a point that stands on its own regardless of how the SLODR argument resolves.
  • A large gap between two index scores is more common at high ability and is not automatically a sign of a problem.
  • Selection at a single high cut score produces a group that is more internally varied than the shared number suggests, which matters for gifted programmes designed as though the admitted students were alike.
  • Comparing two high scores from different batteries is even less safe than comparing two average ones, because the profiles behind them can differ more.

None of that changes what a score is. It remains a position in a reference sample, carrying a published measurement error, and our note on why a reported IQ is a range rather than a point applies at 135 exactly as it does at 100. What the law adds, if it holds at the smaller size the better studies suggest, is a second reason to read the parts: not only is the total imprecise, it is summarising components that have started to go their own ways. The mechanics of that breakdown are set out in our piece on full-scale scores versus index scores.

What the ceiling looks like on a real report

The abstract argument has a concrete counterpart on the page of any score report, and it is worth knowing where to look. On the Wechsler batteries each subtest is reported as a scaled score with a mean of 10 and a standard deviation of 3, running from 1 to 19. Nineteen is not a description of performance; it is the top of the scale. A person who answers every item correctly and a person who answers nearly every item correctly both receive it.

That matters because a full-scale figure built from several subtests sitting at or near their ceiling is averaging numbers that have stopped discriminating. It is why extended norms exist for gifted assessment, allowing scaled scores above 19 to be derived where the standard table has run out, and why an evaluator working at the top of the range will often reach for a battery with more difficult items rather than trusting the total from a general-population instrument. If part of the reason correlations shrink at high ability is that the test has run out of room, then the fix is a harder test rather than a new theory — and separating those two explanations is exactly what makes this literature difficult.

How to read a high result properly

Three habits, in order. Convert the figure to a percentile against the scale it was measured on, because a bare number carries no meaning without its mean and standard deviation — our IQ percentile calculator does that conversion. Then look at the index scores rather than only the total, and note where the profile is uneven. Then treat both the total and each index as bands rather than points. A person described as "135" is more usefully described as someone in roughly the top 1 per cent overall, with particular strengths that the single figure has averaged away. A properly structured assessment reports those parts alongside the total, which is the form worth reading.

Your own number

Where would your own score land?

A score is worth having when you can see the parts it was built from. Take a properly structured assessment, read the index profile as well as the total, and treat both as ranges.

Find your IQ score now!
Secure & encryptedInstant results10–20 minutes

Spearman's qualification has aged better as a caution than as a law. The strong version — that g fades away at the top — is not what the careful studies support. The weak version is both supported and useful: the higher a score sits, the more work a single number is being asked to do, and the more the interesting information has moved into the shape underneath it. That is a good reason to ask what a test measured rather than only what it totalled, which is a reasonable thing to ask of any score at any level.

Common questions

What is Spearman's Law of Diminishing Returns?

It is the observation, made by Charles Spearman in 1927, that different mental tests correlate less strongly with each other among high-ability people than among low-ability people. If it holds, the general factor g accounts for less of the difference between two able people than between two less able people, so a single full-scale score summarises a high performer less completely.

Does g disappear at high IQ?

No. Even the studies most favourable to the law find the general factor still present and still substantial at high ability — it accounts for a smaller share of the variance, not none of it. The defensible version of the claim is that one number tells you proportionally less about a high scorer, not that ability stops being general.

Is Spearman's Law of Diminishing Returns accepted?

It is contested. A 2017 meta-analysis in Intelligence found support across a substantial body of studies, while methodological work has shown that splitting a sample by ability shrinks correlations for statistical reasons alone, so the effect is generally smaller under better controls than the widely cited 1989 figures suggest. Test ceilings at high ability complicate the measurement further.

Should I read my index scores instead of my full-scale IQ?

Read both, and give the index scores more weight the higher the total sits. The full-scale figure pools the most items and is the most reliable single number a battery produces, but it is an average, and at high ability the components it averages tend to be more spread out. An uneven profile above the 98th percentile is common and is not by itself a sign of a problem.

Sources for this story

  1. The Abilities of Man: Their Nature and Measurement, where the ability-differentiation observation is set out — Charles Spearman, Macmillan, 1927
  2. Correlations of mental tests with each other and with cognitive variables are highest for low-IQ groups — Detterman and Daniel, Intelligence, 1989
  3. Spearman's law of diminishing returns: a meta-analysis — Blum and Holling, Intelligence, 2017
  4. Differentiation of cognitive abilities across the life span, on separating age effects from ability effects — Tucker-Drob, Developmental Psychology, 2009
  5. Technical and interpretive manuals for the Wechsler intelligence scales, on index-score structure, subtest intercorrelations and ceiling effects — Pearson

Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.

Share this story

Know someone who keeps seeing this number quoted without the scale it was measured on? Send it to them — it takes one tap.

Filed under#fluid reasoning#scores and scales#study quality#gifted education#percentiles and norms

Read the research.
Then find your own number.

Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.

Start IQ Test
Secure & encryptedInstant results10–20 minutes