Two item types dominate the spatial section of a cognitive battery, and both were built on a finding you can reproduce on yourself: the time it takes to decide whether two shapes match rises in a straight line with how far apart they are rotated.

Open the spatial section of almost any cognitive battery and two item formats account for most of it. One shows a pair of three-dimensional shapes and asks whether they are the same object photographed from different angles. The other shows a square of paper being folded two or three times, a hole punched through the stack, and asks which unfolded sheet results. Neither requires a word of language or a line of arithmetic, which is the point. They are also both descended from specific experiments, and knowing what those experiments found changes how the items look.
Roger Shepard and Jacqueline Metzler published a short paper in Science in 1971 that is among the most reproduced findings in cognitive psychology. They showed people pairs of block figures and asked whether the two were the same shape or mirror images. The interesting measurement was not accuracy but time: the further the second figure was rotated away from the first, the longer people took, and the relationship was close to a straight line.
That linearity is the whole finding. It implies people are not comparing descriptions of the shapes but doing something more like turning an internal image at a steady rate until the two line up — because a rate produces a straight line, and a lookup does not. It also held whether the rotation was in the plane of the picture or in depth, which is why the items on a modern test come in both varieties.
Paper folding items look similar and behave differently. The canonical version is the Paper Folding Test, catalogued as VZ-2 in the ETS Kit of Factor-Referenced Cognitive Tests compiled by Ekstrom and colleagues in 1976, and it is a visualisation task rather than a speeded transformation. You have to hold a sequence of folds, apply the punch to a stack whose layers you cannot see, and then run the folds backwards.
Rotation asks how fast you can turn one image. Folding asks how many states of an object you can hold at once while you change it.
The distinction matters for reading a profile. In the structure most modern batteries follow, both sit under a broad visual-processing ability, separate from verbal comprehension and from quantitative reasoning — but rotation leans on speed and folding leans on working memory, the capacity to hold and manipulate information over seconds. Someone can be reliably good at one and ordinary at the other, and that is not an anomaly in the score sheet. It is the same reason a single total conceals things, which our piece on full-scale scores versus index scores covers in detail.
There is a practical answer with unusually good evidence behind it. Jonathan Wai, David Lubinski and Camilla Benbow reported in the Journal of Educational Psychology in 2009 an analysis of Project Talent — a US study that tested a sample of around 400,000 students in 1960 and followed them for decades. Spatial ability measured in adolescence predicted later entry into and achievement in science, technology, engineering and mathematics, over and above verbal and mathematical scores measured at the same time.
That is a strong result of a kind that is rare: a large sample, a long follow-up, and incremental prediction beyond the measures that usually get all the attention. It is also the argument for why leaving spatial items out of a battery loses information that the verbal and quantitative sections do not recover. The authors' further point was that selection systems built only on verbal and mathematical scores systematically overlook a group who are strong spatially — a finding about admissions and hiring more than about testing.
David Uttal and colleagues published a meta-analysis in Psychological Bulletin in 2013 pooling training studies of spatial skills across age groups and training types. Their conclusion was that spatial ability improves with training, that the improvement is durable, and that it transfers to some untrained spatial tasks. By the standards of cognitive training research — where our note on what brain-training apps actually improve describes a much bleaker picture for transfer — that is an unusually positive result.
It cuts both ways for interpreting a score. It means a low spatial score is not a fixed property and is a reasonable thing to work on, particularly for someone heading into a technical field. It also means a spatial score is more sensitive to prior exposure than, say, a vocabulary score, so a high result from someone who plays a lot of three-dimensional games or has drafting experience carries a different meaning from the same result cold. And any second sitting is affected by the practice effect on top of that — see our note on what a retest actually measures.
For rotation, resolve handedness before attempting to rotate anything: find an asymmetric feature — a limb that turns left rather than right — and check it in both figures. If it disagrees, the pair is a mirror image and no rotation will reconcile them, which saves the several seconds the turning would have cost. For folding, work backwards from the punched sheet rather than forwards through the folds, and count layers: a hole through three layers unfolds to holes in three places, positioned symmetrically about each fold line in reverse order.
It is worth noting that one of these tasks has a strange side career. The spatial item used in the 1993 experiment behind the Mozart effect was a paper folding and cutting task taken from the Stanford-Binet — which is to say that the entire "classical music makes you smarter" industry rests on a short-lived improvement at exactly this kind of puzzle, as our fact check on the Mozart study sets out. The same item type also appears alongside matrix reasoning, and the approach to finding a rule in our note on stating a matrix rule out loud transfers well here: if you cannot say the transformation in a sentence, you have probably not found it. Our tools page collects the calculators and converters that turn any resulting score into something interpretable.
Take the simplest case and the rule falls out of it. Fold a square sheet in half from left to right, then fold it in half again from top to bottom. You now have a quarter-size square made of four layers. Punch a single hole anywhere through the stack and unfold it: there are four holes, one in each quadrant, arranged in mirror image about both fold lines.
That is the whole method, and being able to state it in a sentence is the point. If you are choosing between options by feel rather than by reflecting the hole back through each fold, you are guessing at a task that has an exact procedure — and under time pressure the procedure is faster than the guess.
Spatial reasoning is one of several abilities a full assessment measures separately. Take a properly structured test and read the profile, not just the total.
Find your IQ score now! →What makes these items durable is that they are close to unfakeable and close to language-free. There is no vocabulary to have been taught, no formula to recall, and no way to talk your way to the answer. That is also their limitation: they measure one broad ability well and tell you very little about the others. A spatial score is worth having next to the verbal and quantitative ones, and worth very little on its own — which is true of every single number a test produces, and is the reason the profile exists. A properly structured assessment measures spatial reasoning next to the rest rather than on its own.
A sheet of paper is shown being folded two or three times, a hole is punched through the folded stack, and you choose the pattern of holes on the unfolded sheet. The classic form is the Paper Folding Test, catalogued as VZ-2 in the ETS Kit of Factor-Referenced Cognitive Tests. It measures visualisation — holding an object through a sequence of transformations — rather than speed.
They measure spatial visualisation and the rate at which you can mentally transform an image. Shepard and Metzler showed in Science in 1971 that response time rises roughly linearly with the angle between the two figures, which indicates people solve them by turning an internal representation at a fairly constant rate rather than by comparing descriptions.
Yes, more so than for most cognitive abilities. A 2013 meta-analysis in Psychological Bulletin by Uttal and colleagues found that spatial training produces durable improvement that transfers to some untrained spatial tasks. That also means a spatial score is more sensitive to prior exposure than a vocabulary score, so it should be read with that in mind.
Because spatial ability is distinct from verbal and quantitative ability and predicts outcomes the others miss. Wai, Lubinski and Benbow reported in 2009, using the Project Talent sample of around 400,000 students followed for decades, that adolescent spatial ability predicted later STEM achievement over and above verbal and mathematical scores measured at the same time.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
The test returns a raw count of correct items, not an IQ. Everything the number means arrives later, from a norm table that expires faster than almost any other test's.
Three rows, three columns, two things changing at once. The worked solution is below — and so is an honest account of what solving it does and does not tell you about yourself.

A syllogism gives you two statements and asks what must follow. Decades of research show most people answer with what they already believe instead — and the fix is one simple substitution trick.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →