IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment
Puzzles & Brain TrainingWell established

The Hardest IQ Test Question: What Actually Makes It Hard

There's no single fixed “hardest question” on every IQ test — but the items that sit at the top of a real test's difficulty curve share an exact property, and a landmark cognitive-science study of the Raven's Progressive Matrices shows precisely what it is.

The Hardest IQ Test Question: What Actually Makes It Hard
Illustration generated for IQ Metrics. No photograph is used. IQ Metrics

What you need to know

  • “The hardest IQ question” isn't one fixed riddle. Every properly built test has a difficulty curve, calibrated so its hardest items are solved by only a small share of the norm sample — and that calibration is re-checked every time the test is re-standardized.
  • A landmark 1990 study of the Raven's Advanced Progressive Matrices found the items separating top scorers from everyone else are not the busiest-looking ones. They're the items requiring a “distribution-of-three” rule, which forces the solver to hold all three cells of a row in mind at once instead of comparing two side by side.
  • The single biggest driver of difficulty is how many independent rules an item stacks at once, across how many attributes — easy items vary one thing at a time; the hardest combine several attributes, each governed by its own rule, tracked simultaneously.
  • Viral “impossible IQ riddles” online are almost never drawn from a real, normed test. Genuine test items are kept unpublished specifically because a difficulty rating stops meaning anything once people can look up the answer in advance.

Search “hardest IQ question” and the results are riddles, meme screenshots and forum posts with no test named and no source given. None of it comes from an actual standardized test. Here is the real answer: there is no single hardest question, because difficulty is not a property of one clever riddle — it is a measured, calibrated property of an item inside a specific test, and the research on what actually produces that difficulty is more interesting than any viral riddle.

There is no one “hardest question” — only a difficulty curve

Every well-built cognitive test item has a measured difficulty: the percentage of people in the norm sample who answer it correctly. Test developers deliberately include items across a wide difficulty range, from ones almost everyone gets right to ones only a small fraction solve, because a test made entirely of hard items would fail to distinguish anyone in the middle or lower part of the scale. “The hardest question” on any given test is simply whichever item has the lowest solve rate in its norm sample — and because norm samples get refreshed periodically, that title can shift from one edition of a test to the next.

What a landmark study found inside a real test

The most detailed account of what makes an item hard comes from a 1990 study by Patricia Carpenter, Marcel Just and Peter Shell, published in Psychological Review. They analyzed verbal explanations, eye-movement recordings and error patterns from people solving the Raven's Advanced Progressive Matrices — the classic test of matrix reasoning, where a grid of shapes has one cell missing and the solver has to work out the rule that generates it. Working from Set II of the test, they identified five distinct rule types the items are built from, and showed that item difficulty tracks almost exactly with which rule, or combination of rules, a given item uses.

  • Constant-in-a-row rules (a feature simply stays the same across a row) are the easiest, needing only a quick check.
  • Pairwise-progression rules (a feature changes by a fixed step) are similarly light, solvable by comparing two cells at a time.
  • Distribution-of-three rules are the hard category: the solver must hold all three cells of a row or column in mind at once and reason about which value is missing, a conceptual operation rather than a simple visual comparison.
  • The hardest items combine several of these rules across several attributes of the image at once, each needing to be tracked and satisfied simultaneously.

The hardest items weren't the busiest-looking ones. They were the ones where you have to hold three things in mind at once instead of comparing two.

Carpenter, Just and Shell went further and built two computer simulations of the problem-solving process — one modeling a lower-scoring solver, one modeling a higher-scoring one — to test their theory directly. The model built to solve problems the way high scorers do differed from the low-scorer model mainly in its capacity to manage several sub-goals in working memory at once, rather than in any difference in perceptual sharpness. In plain terms: what separates people at the hard end of a real matrix-reasoning test is less about spotting a pattern and more about how much of the problem they can hold in mind while checking it.

Why stacking rules is the real difficulty engine

Most Raven's items do not use just one rule — they layer two, three or more onto different attributes of the same grid (shape, shading, count, orientation) at once. An easy item might vary only shape, row by row, in a way you can state as a single sentence. A hard item might vary shape one way, a count of internal marks another way, and orientation a third way, all inside the same nine cells — and the answer is only correct if it satisfies every one of those rules simultaneously. We worked through exactly this kind of single-rule item, made explicit step by step, in our piece on a matrix puzzle with one rule you can say out loud; a genuinely hard item is best understood as that same puzzle with two or three more layers stacked on top, each with its own rule to track.

This is also why matrix reasoning specifically, rather than every item format equally, has become the format most associated with “hard” IQ questions in the public imagination. Formats that test a single stored skill — vocabulary, arithmetic fact retrieval — have a difficulty ceiling set mostly by how obscure the content gets. A reasoning item's difficulty ceiling is set by how much working memory and sub-goal tracking the solver can bring to a genuinely novel problem, a steeper and more interesting curve to climb. Our overview of how Raven's Progressive Matrices scores are explained covers how this particular test is scored and interpreted in full. For how any raw score on any test becomes a percentile against a norm sample, our percentile calculator shows the mechanics directly.

“Hard” here is calibrated against a general adult norm sample. An item that is genuinely difficult for the general population may be comparatively easy for someone who works daily with the specific abstraction it uses — part of why a single score from a single sitting is treated as an estimate with a margin of error, not a fixed, context-free fact about a person.

This is also part of why tests need periodic re-norming in the first place. Population-level performance on reasoning-heavy formats like matrix items has shifted across generations — the subject of our piece on why IQ norms expire — so an item's measured difficulty is only accurate against the specific norm sample it was calibrated on, not as some permanent fact about the item itself. A question that was near the top of the difficulty curve fifty years ago is not guaranteed to sit there today.

Why the “impossible IQ riddle” you saw online almost certainly isn't real

Real test items are kept unpublished on purpose. The moment an item's answer is searchable, its measured difficulty stops meaning anything — anyone who has seen it before will solve it regardless of their actual reasoning ability, a problem called item exposure, and it is one of the most closely guarded risks in psychometrics. Test publishers restrict access to items, rotate them, and retire ones that leak. A riddle circulating freely online with no named source, no stated norm sample and no verified solve rate fails the basic definition of a calibrated test item — it is entertainment, not measurement, whatever the caption claims about it.

Your own number

Where would your own score land?

If you want to know where you actually sit on a real, calibrated difficulty scale rather than on one viral riddle, that takes a full test with graded items and a stated norm sample behind it.

Find your IQ score now! →
Secure & encryptedInstant results10–20 minutes

So: there is no single hardest IQ question, viral or otherwise, that means anything outside the specific test it was calibrated on. What decades of research on the format behind most of the internet's “impossible IQ riddles” actually shows is that difficulty comes from stacking independent rules across multiple attributes and forcing the solver to hold more of the problem in mind at once — a finding from real cognitive-science lab work, not a forwarded image with no source. If you want to try graded matrix-style reasoning items yourself under real test conditions, our culture-fair IQ test is built from exactly this item family, and you can see your own result as a percentile on our full IQ test.

Common questions

What is the hardest question on an IQ test?

There isn't one fixed hardest question — difficulty is a measured, calibrated property of each item within a specific test, refreshed whenever the test is re-normed. Research on the Raven's Progressive Matrices, one of the most-studied reasoning tests, found the hardest items are the ones requiring a “distribution-of-three” rule, where the solver must hold three cells in mind at once rather than simply comparing two.

Why are some IQ questions harder than others?

A 1990 study in Psychological Review found difficulty tracks with how many independent rules an item combines and how much working memory it demands. Easy items vary one feature at a time and can be solved by comparing two cells; hard items stack several rules across several attributes at once, requiring the solver to track multiple sub-goals simultaneously.

Are the 'impossible IQ test' riddles that circulate online real test questions?

Almost never. Genuine test items are kept unpublished specifically because a difficulty rating becomes meaningless once an item's answer can be looked up in advance, a problem psychometricians call item exposure. A riddle with no named test, no stated norm sample and no verifiable solve rate does not meet the basic definition of a calibrated item.

Can you get better at the hardest IQ test items with practice?

You can improve meaningfully on that specific item type with practice, which is well documented, but evidence that this transfers into a broader increase in general reasoning ability is much weaker than the evidence for the narrow, practiced gain.

Sources for this story

  1. What one intelligence test measures: a theoretical account of the processing in the Raven Progressive Matrices Test — Psychological Review, 1990 (Carpenter, Just & Shell)
  2. Standards for Educational and Psychological Testing, on item difficulty and item response theory — American Educational Research Association, American Psychological Association and National Council on Measurement in Education
  3. Technical documentation for the Raven's Progressive Matrices, on item calibration and norm samples — test-publisher technical manuals
  4. Research on test-item exposure and security in high-stakes cognitive assessment — psychometric methods literature

Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.

Share this story

Know someone who keeps seeing this number quoted without the scale it was measured on? Send it to them — it takes one tap.

Filed under#raven matrices#fluid reasoning#working memory#puzzles

Read the research.
Then find your own number.

Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.

Start IQ Test →
Secure & encryptedInstant results10–20 minutes