A syllogism gives you two statements and asks what must follow. Decades of research show most people answer with what they already believe instead — and the fix is one simple substitution trick.

Look at these two arguments and decide, quickly, which one is logically valid. First: "All things with fur are warm-blooded. No spiders have fur. Therefore, no spiders are warm-blooded." Second: "All Bloops are Razzles. All Razzles are Cranky. Therefore, all Bloops are Cranky." Most people’s gut says the first one, because the conclusion is true — spiders really are cold-blooded. Their gut is wrong. The first argument is invalid. The second, built entirely from made-up words, is airtight.
A syllogism gives two premises and a conclusion, and asks only one question: given that the premises are true, does the conclusion have to be true as well? That is a narrower question than "is the conclusion true," and it is the one a syllogism question on a test is actually built to check. The form goes back to Aristotle’s Prior Analytics, written in the fourth century BCE, which set out the first systematic rules for when a conclusion validly follows from two categorical statements built from "all," "no" and "some." Modern reasoning tests inherited the structure essentially unchanged.
The spider example fails on form, not content. "All F are W, no S are F" simply does not license any conclusion about whether S is W or is not — the premises leave that question open, and the fact that spiders happen to be cold-blooded is true for reasons the argument never actually establishes. Classical syllogism theory has a plain rule for exactly this shape: no valid conclusion can be drawn when a term that is unrestricted in a premise (here, "warm-blooded") gets treated as fully restricted in the conclusion. The Bloops example, by contrast, is the most basic valid syllogism there is: whatever a Bloop is, if it is a Razzle, and every Razzle is Cranky, a Bloop must be Cranky too. Nonsense content, valid logic.
This is not a trick reserved for confusing test-takers; it is one of the most replicated findings in reasoning research. In a 1983 study published in Memory & Cognition, psychologists Jonathan Evans, Julie Barston and Paul Pollard gave participants syllogisms that crossed logical validity with the believability of the conclusion in every combination — valid-and-believable, valid-and-unbelievable, invalid-and-believable, invalid-and-unbelievable. Judgments on the two combinations where logic and belief point the same way were reliably accurate. Accuracy fell sharply on the two "conflict" combinations, where a valid argument produced an unbelievable conclusion or an invalid one produced a believable conclusion — and people leaned toward belief over logic often enough that the researchers named the pattern belief bias. A hierarchical Bayesian meta-analysis of the accumulated evidence, published decades later in Psychonomic Bulletin & Review, confirmed the effect holds up as one of the sturdier findings in the reasoning literature.
Crucially, the bias is not symmetric. It distorts judgments on invalid syllogisms far more than on valid ones, meaning the dominant failure mode is accepting a believable conclusion that does not actually follow, more than rejecting an unbelievable one that does. Verbal think-aloud studies that had participants reason out loud found the effect is tied to where attention lands first: starting from the conclusion and working backward invites belief bias, while starting from the premises and reasoning forward resists it — a habit worth building deliberately before answering.
Validity is a question about the argument’s shape. Truth is a question about the world. A syllogism question tests whether you can tell which one you are being asked about — and belief bias is what happens when the two get quietly swapped.
The reliable fix is the same move used above: strip out the real-world content and substitute arbitrary terms that preserve the exact structure, then check whether the conclusion still has to hold. If "All A are B, all B are C, therefore all A are C" survives with nonsense words standing in for A, B and C, it survives with any words at all — that particular form, traditionally labeled Barbara, is always valid. Classical logic actually counts this precisely: with four ways to phrase each premise and the conclusion, and four possible positions for the shared middle term, there are exactly 256 distinct categorical-syllogism forms — and only 24 of them are valid under any circumstances. Recognizing a handful of the common ones, Barbara among them, covers most of what a reasoning test will actually ask. A second rule catches a common trap without any substitution at all: two negative premises — two statements built on "no" or "not" — never yield a valid conclusion of any kind, regardless of subject matter, because two negatives only ever tell you what is excluded, never what must be included.
One more pair to test the rule before moving on. "No reptiles are mammals. All snakes are reptiles. Therefore, no snakes are mammals" — swap in nonsense terms and the structure survives (no A are B, all C are A, therefore no C are B is a valid form), so this one is sound, and it also happens to be true. Compare: "Some birds cannot fly. All penguins are birds. Therefore, some penguins cannot fly." That conclusion is also true, but the argument is not valid — "some birds cannot fly" does not guarantee penguins specifically are among them, and substituting nonsense terms makes the gap obvious immediately. A true conclusion and a valid argument are two different achievements, and only one of these two pairs delivers both.
Syllogism-style items are not confined to puzzle books. The LSAT’s Logical Reasoning section, the Watson-Glaser Critical Thinking Appraisal used in law and corporate hiring, and employer cognitive-ability batteries such as the CCAT and SHL’s verbal reasoning tests all include deduction items built on the same categorical structure. The reason is consistent across all of them: a syllogism isolates the ability to reason from given information, cleanly separated from vocabulary size, general knowledge or subject-matter expertise. That is the same property our piece on analogy questions on IQ tests covers for a different question type, and it is why both formats keep reappearing across otherwise very different tests, including the general cognitive-ability tests used in hiring.
Curious how you handle deduction under time pressure? Our full IQ assessment includes reasoning items and reports a percentile against a real reference population, not just a raw right-or-wrong count.
Find your IQ score now! →The next syllogism you meet — on a test, in an argument, in a headline — is easier to check than it looks. Ignore whether the conclusion sounds right. Ask only whether it had to follow. Our explainer on how IQ tests are built covers why that particular skill, reasoning from given premises rather than from prior belief, is one of the most consistently measured abilities on record. Take the full test if you want to see where you land on it.
A short argument made of two premises and a conclusion, where the question is whether the conclusion necessarily follows from the premises — not whether the statements are true in the real world.
Research on "belief bias" shows people tend to judge a syllogism by whether its conclusion sounds true rather than by whether it logically follows, and this bias is strongest on invalid arguments with believable conclusions — the exact combination most likely to be marked wrong.
Replace the real-world nouns with arbitrary placeholder words, keeping the quantifiers ("all," "no," "some") unchanged, and check whether the conclusion still has to be true. If it does not, the argument is invalid regardless of the original topic.
The LSAT’s Logical Reasoning section, the Watson-Glaser Critical Thinking Appraisal, and employer cognitive-ability tests including the CCAT and SHL’s verbal reasoning batteries all include categorical-syllogism-style deduction items.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
Every analogy item asks the same thing twice. Name the relation between the first pair in a single sentence, then find the pair that takes the same sentence — and the plausible-looking wrong answers stop being tempting.
Aptitude tests are among the better-studied hiring tools, and for decades the headline validity figures were quoted with more confidence than the corrections behind them deserved. A 2022 reanalysis pulled those numbers down.

Four items share a rule and one breaks it, but only if that rule is the only sentence that fits. When a second, equally valid rule points at a different item, the puzzle is testing which distractor you noticed first, not how well you reason.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →