Cognitive tests correlate less strongly with each other among high scorers than among low scorers. If that holds, a single full-scale figure carries less information the higher it goes — and the evidence for it is real but smaller and shakier than the idea's popularity suggests.

Two people take the same battery. One scores 85 on a mean-100, standard-deviation-15 scale; the other scores 135. Both walk away with a single number that is supposed to describe how they did across every part of the test. There is a century-old argument in psychometrics that those two numbers are not doing equally good jobs — that the higher one is a worse summary of the person it describes. The claim is called Spearman's Law of Diminishing Returns, and it is one of the few ideas in intelligence research where the popular version and the evidence have drifted noticeably apart.
Charles Spearman is the reason anyone talks about a general factor at all. His central finding, first set out in 1904 and developed through the 1920s, was that people who do well on one kind of mental test tend to do well on the others — vocabulary, arithmetic, spatial puzzles, memory span. Every pair of tests correlates positively, a pattern known as the positive manifold, and the shared variance is what the letter g refers to. Our explainer on what the g factor actually is covers that ground in full.
In his 1927 book The Abilities of Man, Spearman added a qualification that is far less often quoted. The strength of those correlations, he noted, was not constant across the range of ability. Among people who did poorly, the tests hung together tightly. Among people who did well, they came apart. He described the general factor as behaving like a resource that mattered enormously when it was scarce and mattered progressively less once there was plenty of it — hence the name later attached to the idea.
A correlation between two subtests describes how well one predicts the other. If verbal comprehension and matrix reasoning correlate at 0.55 in a group, knowing one score tells you a fair amount about the other. If they correlate at 0.30, it tells you much less. Spearman's claim is that the first figure describes low scorers and the second describes high scorers — and if that is right, several things follow:
At the middle of the distribution one number is a fair summary of a person. At the top it is increasingly a summary of an average of things that have started to diverge.
The modern literature starts with a 1989 study by Douglas Detterman and Mark Daniel, who split standardisation samples by ability and compared the subtest correlation matrices. In the lowest-ability groups the average intercorrelation was in the region of twice what it was in the highest-ability groups. That result is the one usually cited, and it is a large effect — large enough that if it were the whole story, test manuals would need to say so on the first page.
It is not the whole story. Splitting a sample by ability and then measuring correlations inside each slice introduces a well-known statistical artefact: restricting the range of a variable reduces its correlations mechanically, whether or not anything about the underlying structure has changed. Much of the work since has been an argument about how to separate a real change in structure from that artefact. Approaches that model the change continuously rather than by cutting the sample into groups — moderated factor analysis, and methods that estimate the factor structure locally along the ability range — generally do find the effect in Spearman's direction, but smaller than the 1989 figures imply. A 2017 meta-analysis by Blum and Holling in Intelligence concluded that support for the law exists across a substantial body of studies while being moderated by sample and method. Other analyses, including work on how abilities differentiate across the lifespan, have found age effects considerably clearer than ability effects.
The practical reading is narrower than the idea's reputation. Nobody serious claims g stops existing at high ability, or that a 135 is uninformative. The claim is about how much of the story one number tells, and it bears on a specific set of decisions:
None of that changes what a score is. It remains a position in a reference sample, carrying a published measurement error, and our note on why a reported IQ is a range rather than a point applies at 135 exactly as it does at 100. What the law adds, if it holds at the smaller size the better studies suggest, is a second reason to read the parts: not only is the total imprecise, it is summarising components that have started to go their own ways. The mechanics of that breakdown are set out in our piece on full-scale scores versus index scores.
The abstract argument has a concrete counterpart on the page of any score report, and it is worth knowing where to look. On the Wechsler batteries each subtest is reported as a scaled score with a mean of 10 and a standard deviation of 3, running from 1 to 19. Nineteen is not a description of performance; it is the top of the scale. A person who answers every item correctly and a person who answers nearly every item correctly both receive it.
That matters because a full-scale figure built from several subtests sitting at or near their ceiling is averaging numbers that have stopped discriminating. It is why extended norms exist for gifted assessment, allowing scaled scores above 19 to be derived where the standard table has run out, and why an evaluator working at the top of the range will often reach for a battery with more difficult items rather than trusting the total from a general-population instrument. If part of the reason correlations shrink at high ability is that the test has run out of room, then the fix is a harder test rather than a new theory — and separating those two explanations is exactly what makes this literature difficult.
Three habits, in order. Convert the figure to a percentile against the scale it was measured on, because a bare number carries no meaning without its mean and standard deviation — our IQ percentile calculator does that conversion. Then look at the index scores rather than only the total, and note where the profile is uneven. Then treat both the total and each index as bands rather than points. A person described as "135" is more usefully described as someone in roughly the top 1 per cent overall, with particular strengths that the single figure has averaged away. A properly structured assessment reports those parts alongside the total, which is the form worth reading.
A score is worth having when you can see the parts it was built from. Take a properly structured assessment, read the index profile as well as the total, and treat both as ranges.
Find your IQ score now! →Spearman's qualification has aged better as a caution than as a law. The strong version — that g fades away at the top — is not what the careful studies support. The weak version is both supported and useful: the higher a score sits, the more work a single number is being asked to do, and the more the interesting information has moved into the shape underneath it. That is a good reason to ask what a test measured rather than only what it totalled, which is a reasonable thing to ask of any score at any level.
It is the observation, made by Charles Spearman in 1927, that different mental tests correlate less strongly with each other among high-ability people than among low-ability people. If it holds, the general factor g accounts for less of the difference between two able people than between two less able people, so a single full-scale score summarises a high performer less completely.
No. Even the studies most favourable to the law find the general factor still present and still substantial at high ability — it accounts for a smaller share of the variance, not none of it. The defensible version of the claim is that one number tells you proportionally less about a high scorer, not that ability stops being general.
It is contested. A 2017 meta-analysis in Intelligence found support across a substantial body of studies, while methodological work has shown that splitting a sample by ability shrinks correlations for statistical reasons alone, so the effect is generally smaller under better controls than the widely cited 1989 figures suggest. Test ceilings at high ability complicate the measurement further.
Read both, and give the index scores more weight the higher the total sits. The full-scale figure pools the most items and is the most reliable single number a battery produces, but it is an average, and at high ability the components it averages tend to be more spread out. An uneven profile above the 98th percentile is common and is not by itself a sign of a problem.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
Scores on tests that look nothing alike still correlate positively with one another. g is the factor pulled out of that pattern — which is what makes a single summary number defensible, and also what makes it far less than the whole story.
A Full-Scale IQ is an average of several narrower index scores, and when those indexes disagree by enough, the average describes none of them well.

A full-scale IQ is not one thing a test measures directly. It is a composite built from several index scores, and the theory that decides which ones is called Cattell-Horn-Carroll — CHC for short.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →