IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment
Resources

How IQ Tests
Work.

Diagram of a hierarchical model of cognitive ability, with general ability g at the apex feeding into broad and narrow abilities
Mean 100 · SD 15The standard scale
Since 1904Spearman’s g factor

How can a handful of shapes, patterns and sequences say anything about how a mind works? The answer rests on a century of psychometric research — and on one stubborn statistical finding that keeps turning up no matter how the questions are written.

4.8 out of 5 Excellent 4.8/51 million+ test takersIIF certified
4.8 / 5Average rating
1,000,000+Tests taken
10–20 minInstant results
IIF CertifiedEndorsed by psychologists
The architecture of ability

One finding that refused to go away

By the IQ Metrics Research Team · A plain-language psychometric overview

In 1904, the psychologist Charles Spearman noticed something odd in schoolchildren’s marks. Performance across quite different subjects was positively correlated: pupils who did well at one tended to do well at others, even when the tasks appeared to share nothing obvious. That pattern — later nicknamed the positive manifold — has been replicated with remarkable consistency ever since.

Spearman’s explanation was that every mental task draws on two things: a general factor (g) common to all cognitive work, and a specific factor (s) unique to that particular task. Vocabulary and mental rotation look nothing alike, but if both tap the same general capacity, scores on them will move together. That is what a modern IQ test is built to estimate — not any single skill, but the shared variance running underneath many of them.

Later work refined the picture rather than replacing it. Drawing on John Carroll’s large-scale 1993 reanalysis of hundreds of datasets, and on the earlier Cattell–Horn tradition, the field converged on the Cattell–Horn–Carroll (CHC) model: a three-level hierarchy with g at the top, a set of broad abilities beneath it, and dozens of narrow abilities below those. It is the framework most contemporary cognitive batteries are organised around.

STRATUM III General ability STRATUM II Broad abilities STRATUM I Narrow abilities g Fluid (Gf) Crystallized (Gc) Visual (Gv) Speed (Gs) InductionReasoning LanguageLexical VisualizeSpatial PerceptualReaction examples only — CHC taxonomies list dozens more narrow abilities per domain

Figure 1. The CHC hierarchy, simplified. General ability sits at the apex; broad abilities such as fluid reasoning and comprehension-knowledge sit beneath it; and each of those subdivides into many narrow abilities. Only four broad abilities are shown here — published CHC taxonomies list considerably more.

Cattell’s distinction

Reasoning it out vs. already knowing

The single most useful split in the whole model — and the reason non-verbal tests look the way they do.

Working with John Horn, Raymond Cattell argued that general ability is not one undifferentiated thing but divides into two broad kinds. Fluid intelligence (Gf) is the capacity to reason your way through an unfamiliar problem: spotting a pattern, inferring a rule, holding several possibilities in mind at once — with little help from what you happen to have learned. Crystallized intelligence (Gc) is the accumulated result of learning: vocabulary, factual knowledge, and comprehension built up through education and experience.

Fluid and crystallized intelligence compared
Dimension Fluid intelligence (Gf) Crystallized intelligence (Gc)
What it captures Reasoning about genuinely novel problems Knowledge and comprehension already acquired
Typical items Visual matrices, series completion, abstract analogies Vocabulary, verbal comprehension, general knowledge
Prior learning Minimised by design Central — it is learned material
Across adulthood Tends to peak earlier, then decline gradually Tends to hold steady or keep rising into later life

Lifespan patterns describe general population trends, not any individual’s trajectory.

This is exactly why non-verbal, matrix-style tests carry so much weight in cross-cultural assessment. A vocabulary question inevitably measures which language you speak and how much schooling you have had alongside your reasoning. A visual pattern with no words in it leans much more heavily on Gf. That reduction in cultural loading is real and useful — but “culture-fair” means reduced, not eliminated. Familiarity with abstract diagrams and formal test conventions still varies between people.

Inside a single item

How a matrix question actually works

A 3×3 matrix looks like a puzzle. Structurally, it is a test of whether you can extract rules from evidence and apply them to a case you have not seen.

ACROSS → ORIENTATION ROTATES DOWN ↓ SHADING CHANGES ? ANSWER

Figure 2. A simplified two-rule matrix. Orientation advances 90° across each row; shading runs solid → hatched → outline down each column. The missing cell must satisfy both at once, giving an unfilled triangle at 180°. Published test items are not this tidy — this is a teaching example, not a real item.

1

Read across the rows

Compare left to right and one property changes systematically. Here the triangle’s orientation advances by 90° at each step.

2

Read down the columns

A second, independent property varies vertically — solid, then hatched, then outline. Neither rule explains the grid on its own.

3

Apply both at once

The answer has to satisfy every rule simultaneously. Harder items add a third or fourth rule, and the load of holding them all in mind is where difficulty really comes from.

From answers to a number

How raw performance becomes an IQ score

The number on your report is not a mark out of a hundred. It is a position on a distribution.

01

Score the items

Responses are marked and weighted, since a hard item that few people solve carries more information than an easy one nearly everyone gets right.

02

Aggregate the domains

Performance is summed within each reasoning domain and then combined, so the composite reflects the shared reasoning ability rather than one strong area.

03

Standardise against age

The composite is compared with the distribution for the same age group and converted onto the familiar scale — average 100, standard deviation 15.

That third step is the one people most often misread. Because scoring is age-normed, an identical set of correct answers can map to different IQ scores at different ages — the comparison group changes. And because the scale is built from a distribution rather than a percentage, the same gap in points means very different things depending on where it sits: the difference between 98 and 105 is crowded with people, while the same seven points out at the tail separates far rarer performances.

7085100115130AVERAGEabout 68% of people

Figure 3. The standard distribution used throughout cognitive assessment: an average of 100 with a standard deviation of 15. Roughly 68% of people score between 85 and 115 — the shaded band. Scores further from the centre are progressively rarer, which is why a few points near the middle mean far less than the same few points out at the edges.

This is also why a single score should be read with its margin of error in mind. Every standardised cognitive measure carries some measurement error, so a result is better understood as a band around a value than as a precise point — a normal property of psychometric testing rather than a flaw peculiar to any one test.

Questions

Frequently asked

Why do IQ tests use so many different question types?

Each item type draws on a different specific ability as well as on general reasoning. By sampling across verbal, numerical, spatial and matrix items, a test averages out the quirks of any single format — so what the different sections share gives a more stable estimate of general cognitive ability than any one item type could on its own. Our IQ Tests hub breaks down what each format actually measures.

What is the difference between an IQ test and an aptitude test?

Aptitude tests are built to predict performance in a particular setting — a job role, a course of study — and are often tied closely to relevant content. IQ tests aim at something broader: general reasoning capacity across domains, expressed relative to an age group. The two overlap, because general reasoning predicts a wide range of outcomes, but their purposes differ. Our IQ Test Guide covers how to read a general-reasoning score correctly.

Are visual matrix tests genuinely culture-fair?

They substantially reduce dependence on language and on culturally specific knowledge, which is why they are widely used for cross-cultural assessment. But no test is entirely free of cultural influence — familiarity with abstract diagrams, formal schooling and test-taking conventions all still play a part. Read “culture-fair” as reduced cultural loading, not none. Our Culture-Fair IQ Test is built around exactly this format.

Does a higher IQ score mean I answered a higher percentage correctly?

No — an IQ score is not a percentage. It is a standardised rank describing where your performance sits relative to others of a similar age, on a scale with an average of 100 and a standard deviation of 15. The same number of correct answers can produce different scores in different age groups. Our IQ Score Converter shows how raw performance maps onto that scale.

Why do some IQ tests use a strict time limit while others don’t?

Timed formats add processing speed (Gs) into the composite score, since how quickly you can work through items reliably is itself a measured ability. Untimed formats isolate reasoning power from speed, which matters for people who process carefully rather than fast. Neither is more valid than the other — they measure a related but distinct mix. Our Timed IQ Test is built specifically around this speed component.

How different are fluid and crystallized intelligence in day-to-day terms?

Fluid intelligence is what you use to work out a problem you’ve never seen before with no relevant prior knowledge; crystallized intelligence is what you use to recall a fact, a word, or a learned procedure. Most real tasks blend both, which is why full IQ tests weigh contributions from each rather than measuring just one. See how each changes across a lifetime on our Can You Improve Your IQ page.

How exactly is my raw score converted into a percentile?

Your raw number of correct answers is first compared against the norm sample for your age group, converted to a standardised score centred on 100, and then mapped onto the normal distribution to produce a percentile — the share of that norm group your score is higher than. Our IQ Percentile Calculator walks through this conversion directly.

Why does the bell curve matter for interpreting my score?

IQ scores are constructed to follow a normal distribution, so the same numeric gap means very different things depending on where it falls. A move from 100 to 105 shifts your percentile only slightly, while a move from 130 to 135 shifts it much less than that gap size would suggest near the middle. Our IQ Bell Curve page shows this relationship visually.

Does comparing my score to a different age group change the result?

Yes, substantially. Because scores are normed against an age-matched sample, the same raw performance produces a different standardised score when compared against a younger or older reference group, since typical fluid and crystallized ability levels shift across the lifespan. Our IQ and Age tool breaks down how that adjustment works.

What does a verbal reasoning section actually measure?

Verbal sections typically test analogies, classification, vocabulary and comprehension — tapping largely into crystallized intelligence (Gc), since they depend on language knowledge you’ve already acquired rather than novel abstract reasoning. Our Verbal IQ Test isolates this domain specifically.

What does a numerical reasoning section actually measure?

Numerical sections cover number series, quantitative comparison, data interpretation and estimation — drawing on quantitative reasoning, a narrow ability under fluid intelligence, rather than on taught arithmetic procedures. Our Numerical IQ Test is built around this distinction, not a maths curriculum.

How much should I trust the result from a single test session?

Any single session carries some measurement error from factors like fatigue, stress or an unfamiliar format, so a single score is best read as an estimate with a margin around it rather than an exact number. Our IQ Test Accuracy page covers how large that margin typically is and when a retest is worth doing.

See the rules for yourself.

Work through visual matrix questions with the transformation explained step by step — or take the scored, language-free assessment built on the same principles.

Secure & encrypted Free to practise 10–20 minutes