How can a handful of shapes, patterns and sequences say anything about how a mind works? The answer rests on a century of psychometric research — and on one stubborn statistical finding that keeps turning up no matter how the questions are written.
Excellent 4.8/51 million+ test takersIIF certified
By the IQ Metrics Research Team · A plain-language psychometric overview
In 1904, the psychologist Charles Spearman noticed something odd in schoolchildren’s marks. Performance across quite different subjects was positively correlated: pupils who did well at one tended to do well at others, even when the tasks appeared to share nothing obvious. That pattern — later nicknamed the positive manifold — has been replicated with remarkable consistency ever since.
Spearman’s explanation was that every mental task draws on two things: a general factor (g) common to all cognitive work, and a specific factor (s) unique to that particular task. Vocabulary and mental rotation look nothing alike, but if both tap the same general capacity, scores on them will move together. That is what a modern IQ test is built to estimate — not any single skill, but the shared variance running underneath many of them.
Later work refined the picture rather than replacing it. Drawing on John Carroll’s large-scale 1993 reanalysis of hundreds of datasets, and on the earlier Cattell–Horn tradition, the field converged on the Cattell–Horn–Carroll (CHC) model: a three-level hierarchy with g at the top, a set of broad abilities beneath it, and dozens of narrow abilities below those. It is the framework most contemporary cognitive batteries are organised around.
Figure 1. The CHC hierarchy, simplified. General ability sits at the apex; broad abilities such as fluid reasoning and comprehension-knowledge sit beneath it; and each of those subdivides into many narrow abilities. Only four broad abilities are shown here — published CHC taxonomies list considerably more.
The single most useful split in the whole model — and the reason non-verbal tests look the way they do.
Working with John Horn, Raymond Cattell argued that general ability is not one undifferentiated thing but divides into two broad kinds. Fluid intelligence (Gf) is the capacity to reason your way through an unfamiliar problem: spotting a pattern, inferring a rule, holding several possibilities in mind at once — with little help from what you happen to have learned. Crystallized intelligence (Gc) is the accumulated result of learning: vocabulary, factual knowledge, and comprehension built up through education and experience.
| Dimension | Fluid intelligence (Gf) | Crystallized intelligence (Gc) |
|---|---|---|
| What it captures | Reasoning about genuinely novel problems | Knowledge and comprehension already acquired |
| Typical items | Visual matrices, series completion, abstract analogies | Vocabulary, verbal comprehension, general knowledge |
| Prior learning | Minimised by design | Central — it is learned material |
| Across adulthood | Tends to peak earlier, then decline gradually | Tends to hold steady or keep rising into later life |
Lifespan patterns describe general population trends, not any individual’s trajectory.
This is exactly why non-verbal, matrix-style tests carry so much weight in cross-cultural assessment. A vocabulary question inevitably measures which language you speak and how much schooling you have had alongside your reasoning. A visual pattern with no words in it leans much more heavily on Gf. That reduction in cultural loading is real and useful — but “culture-fair” means reduced, not eliminated. Familiarity with abstract diagrams and formal test conventions still varies between people.
A 3×3 matrix looks like a puzzle. Structurally, it is a test of whether you can extract rules from evidence and apply them to a case you have not seen.
Figure 2. A simplified two-rule matrix. Orientation advances 90° across each row; shading runs solid → hatched → outline down each column. The missing cell must satisfy both at once, giving an unfilled triangle at 180°. Published test items are not this tidy — this is a teaching example, not a real item.
Compare left to right and one property changes systematically. Here the triangle’s orientation advances by 90° at each step.
A second, independent property varies vertically — solid, then hatched, then outline. Neither rule explains the grid on its own.
The answer has to satisfy every rule simultaneously. Harder items add a third or fourth rule, and the load of holding them all in mind is where difficulty really comes from.
The number on your report is not a mark out of a hundred. It is a position on a distribution.
Responses are marked and weighted, since a hard item that few people solve carries more information than an easy one nearly everyone gets right.
Performance is summed within each reasoning domain and then combined, so the composite reflects the shared reasoning ability rather than one strong area.
The composite is compared with the distribution for the same age group and converted onto the familiar scale — average 100, standard deviation 15.
That third step is the one people most often misread. Because scoring is age-normed, an identical set of correct answers can map to different IQ scores at different ages — the comparison group changes. And because the scale is built from a distribution rather than a percentage, the same gap in points means very different things depending on where it sits: the difference between 98 and 105 is crowded with people, while the same seven points out at the tail separates far rarer performances.
Figure 3. The standard distribution used throughout cognitive assessment: an average of 100 with a standard deviation of 15. Roughly 68% of people score between 85 and 115 — the shaded band. Scores further from the centre are progressively rarer, which is why a few points near the middle mean far less than the same few points out at the edges.
This is also why a single score should be read with its margin of error in mind. Every standardised cognitive measure carries some measurement error, so a result is better understood as a band around a value than as a precise point — a normal property of psychometric testing rather than a flaw peculiar to any one test.
Each item type draws on a different specific ability as well as on general reasoning. By sampling across verbal, numerical, spatial and matrix items, a test averages out the quirks of any single format — so what the different sections share gives a more stable estimate of general cognitive ability than any one item type could on its own. Our IQ Tests hub breaks down what each format actually measures.
Aptitude tests are built to predict performance in a particular setting — a job role, a course of study — and are often tied closely to relevant content. IQ tests aim at something broader: general reasoning capacity across domains, expressed relative to an age group. The two overlap, because general reasoning predicts a wide range of outcomes, but their purposes differ. Our IQ Test Guide covers how to read a general-reasoning score correctly.
They substantially reduce dependence on language and on culturally specific knowledge, which is why they are widely used for cross-cultural assessment. But no test is entirely free of cultural influence — familiarity with abstract diagrams, formal schooling and test-taking conventions all still play a part. Read “culture-fair” as reduced cultural loading, not none. Our Culture-Fair IQ Test is built around exactly this format.
No — an IQ score is not a percentage. It is a standardised rank describing where your performance sits relative to others of a similar age, on a scale with an average of 100 and a standard deviation of 15. The same number of correct answers can produce different scores in different age groups. Our IQ Score Converter shows how raw performance maps onto that scale.
Timed formats add processing speed (Gs) into the composite score, since how quickly you can work through items reliably is itself a measured ability. Untimed formats isolate reasoning power from speed, which matters for people who process carefully rather than fast. Neither is more valid than the other — they measure a related but distinct mix. Our Timed IQ Test is built specifically around this speed component.
Fluid intelligence is what you use to work out a problem you’ve never seen before with no relevant prior knowledge; crystallized intelligence is what you use to recall a fact, a word, or a learned procedure. Most real tasks blend both, which is why full IQ tests weigh contributions from each rather than measuring just one. See how each changes across a lifetime on our Can You Improve Your IQ page.
Your raw number of correct answers is first compared against the norm sample for your age group, converted to a standardised score centred on 100, and then mapped onto the normal distribution to produce a percentile — the share of that norm group your score is higher than. Our IQ Percentile Calculator walks through this conversion directly.
IQ scores are constructed to follow a normal distribution, so the same numeric gap means very different things depending on where it falls. A move from 100 to 105 shifts your percentile only slightly, while a move from 130 to 135 shifts it much less than that gap size would suggest near the middle. Our IQ Bell Curve page shows this relationship visually.
Yes, substantially. Because scores are normed against an age-matched sample, the same raw performance produces a different standardised score when compared against a younger or older reference group, since typical fluid and crystallized ability levels shift across the lifespan. Our IQ and Age tool breaks down how that adjustment works.
Verbal sections typically test analogies, classification, vocabulary and comprehension — tapping largely into crystallized intelligence (Gc), since they depend on language knowledge you’ve already acquired rather than novel abstract reasoning. Our Verbal IQ Test isolates this domain specifically.
Numerical sections cover number series, quantitative comparison, data interpretation and estimation — drawing on quantitative reasoning, a narrow ability under fluid intelligence, rather than on taught arithmetic procedures. Our Numerical IQ Test is built around this distinction, not a maths curriculum.
Any single session carries some measurement error from factors like fatigue, stress or an unfamiliar format, so a single score is best read as an estimate with a margin around it rather than an exact number. Our IQ Test Accuracy page covers how large that margin typically is and when a retest is worth doing.
Work through visual matrix questions with the transformation explained step by step — or take the scored, language-free assessment built on the same principles.