Cognitive Reflection Test: What the Bat-and-Ball Score Says About IQ
Across 3,428 people the three-question Cognitive Reflection Test averaged 1.24 correct, and only 17% got all three. Its link to IQ is moderate, not one-to-one, and the questions are now so well known that familiarity shifts scores.

What you need to know
- The Cognitive Reflection Test (CRT) is three short puzzles, each with a tempting wrong answer. Shane Frederick's 2005 paper pooled 3,428 respondents from 35 studies: the mean was 1.24 out of 3, 33% scored zero and 17% scored three. The bat-and-ball answer is 5 cents, not 10.
- A 2022 meta-analysis of 49 samples (14,905 people) found the CRT correlates r = .38 with cognitive intelligence, .61 after statistical corrections. It correlates more strongly with numeracy (.79 corrected), so it overlaps with number skill as much as with IQ.
- There is no pass mark and no official IQ conversion. Our arithmetic, which assumes a normal curve, suggests a perfect 3 out of 3 fits an average IQ of roughly 109 to 114 and a zero fits roughly 90 to 94, with wide spread around both.
- Familiarity matters. In one study 51.4% of respondents had already seen an item and scored 2.36 against 1.48 for those who had not. In a four-second first-response experiment, 67.2% of those who finally answered correctly had been right at the first answer.
The Cognitive Reflection Test (CRT) is a three-question puzzle test designed by Shane Frederick of MIT, and it is best known for the bat-and-ball problem, whose correct answer is 5 cents and not the 10 cents most people blurt out. Across the 3,428 respondents in Frederick's 2005 paper the average score was 1.24 out of 3, and only 17% answered all three correctly. The CRT is not an IQ test and has no IQ conversion table, but it does correlate with intelligence. This guide gives the three questions and their answers, what scores look like by university, how strongly the CRT tracks IQ, how much having seen it before matters, and how chatbots fare on it.
What are the three Cognitive Reflection Test questions and answers?
Each of the three items has an answer that springs to mind at once and is wrong, and a correct answer that takes a few seconds of checking. Frederick built the test that way on purpose: the score measures whether you stop to check the first answer, not whether you can do the arithmetic, which is easy once you slow down.
- Bat and ball: a bat and a ball cost $1.10 together and the bat costs $1.00 more than the ball. The ball costs 5 cents, not the intuitive 10
- Machines and widgets: if 5 machines take 5 minutes to make 5 widgets, 100 machines take 5 minutes to make 100, not the intuitive 100 minutes
- Lily pads: if a patch of lily pads doubles each day and takes 48 days to cover a lake, it covers half the lake on day 47, not the intuitive day 24
The way to the right answer is to write the relationship down. For the ball, call its price x, so the bat is x plus $1.00 and the total is 2x plus $1.00 = $1.10, which gives x = 5 cents; at 10 cents the bat would be $1.10 and the total $1.20. For the machines, each machine makes one widget in 5 minutes, so 100 machines make 100 widgets in the same 5 minutes. For the lake, work backwards: a patch that doubles daily is half the size one day before it is full. Our guide to what IQ test items measure shows what standard batteries ask for instead, and why trick riddles are not IQ test items explains why the CRT counts as a research instrument measuring a narrower trait than general fluid intelligence.
What is the average Cognitive Reflection Test score?
Frederick's Table 1, in the Journal of Economic Perspectives (2005, volume 19, issue 4, pages 25 to 42), pools 3,428 respondents across 35 studies. The overall mean is 1.24: 33% scored 0, 28% scored 1, 23% scored 2 and 17% scored 3. It also reports means for named university samples, which show how much the comparison group changes what a good score is.
- MIT (61 students): mean 2.18; 48% scored 3 and 7% scored 0
- Princeton (121): mean 1.63; 26% scored 3 and 18% scored 0
- Carnegie Mellon (746): mean 1.51; 25% scored 3 and 25% scored 0
- Harvard (51): mean 1.43; 20% scored 3 and 20% scored 0
- All 3,428 respondents: mean 1.24; 17% scored 3 and 33% scored 0
Two things follow. First, there is no official pass mark: these are convenience samples collected by the researcher, not a population norm, so a good score is only a score high against a stated group. Against the pooled sample, 3 out of 3 puts you in the top 17%, and 2 or 3 in the top 40%. Second, even at the most selective schools the test humbles people: 7% of the MIT sample and one in five Harvard students scored zero. The scale is also very coarse. With four possible scores, 61% of the pooled sample landed on 0 or 1, so the CRT cannot separate people the way a 40-item test can; our guide to the IQ score margin of error shows how little precision a short scale carries.
Does the Cognitive Reflection Test predict IQ?
It predicts it moderately. In Frederick's own data the CRT correlated .44 with SAT scores (434 students), .46 with ACT scores (667) and .43 with the Wonderlic cognitive test (921); our guide to Wonderlic scores and IQ covers that instrument. Self-reported SAT and ACT scores were used, and the link to the Need for Cognition questionnaire, a measure of how much people enjoy thinking, was weaker at .22. Toplak, West and Stanovich (2011) found the CRT correlated about .40 with a composite of cognitive-ability measures.
The best summary is the 2022 meta-analysis by Otero, Salgado and Moscoso in the journal Intelligence. Across 49 samples and 14,905 people the CRT correlated r = .38 with cognitive intelligence, and .61 (confidence interval .56 to .66) after correcting for measurement error and similar limits. The surprise is the second result: with numeracy, the ability to work with numbers, the corrected correlation was .79 across 44 samples and 20,307 people, higher than with intelligence. The original three-item version alone gave .64 with intelligence. So the test is partly a numeracy test, which fits the fact that the traps are all numerical.
What do those correlations mean for one person? A correlation of .38 means about 14% of the variation in CRT scores is shared with intelligence; .61 means about 37%. We worked out an illustration, assuming that CRT performance and IQ follow a bivariate normal distribution and treating Frederick's pooled distribution as the reference group, which it is not. The 17% who scored 3 would then average an IQ of about 109 at r = .38 and about 114 at .61, and the 33% who scored zero would average about 94 at r = .38 and about 90 at .61. Those are group averages with a wide spread, so plenty of perfect scorers sit below 100 and plenty of zero scorers above it. Our percentile calculator, the IQ bell curve and the guide to IQ z-scores show the arithmetic behind a shift of this size.
A perfect 3 out of 3 on the Cognitive Reflection Test fits an average IQ of roughly 109 to 114: above most people, nowhere near a genius label.
Does seeing the questions before change your score?
Yes, and by a lot. Haigh (2016) found that 51.4% of 142 respondents had already seen at least one CRT item, and that they scored 2.36 on average against 1.48 for those who had not. Stieger and Reips (2016), with 2,272 respondents, reported that 44% had prior experience and an effect size of d = 0.41. Bialek and Pennycook (2018) then ran six studies with about 2,500 people and concluded that the test remains a valid predictor of reasoning even after repeat exposure. The score has become less clean, while the link with other reasoning measures has largely survived.
Test designers answered with longer versions. Thomson and Oppenheimer (2016) published a four-item CRT-2 with new problems; it correlates .511 with the original. The same exposure problem hits every well-known test, as our guide to the practice effect on cognitive ability tests explains. The broader point for anyone who has seen the bat-and-ball puzzle before is that a correct answer now partly measures memory.
Is the right answer really a product of slow thinking?
The popular story says fast intuition gives 10 cents and slow reflection fixes it to 5. Bago and De Neys (2019) tested that on the bat-and-ball problem itself. In their first study participants had to give a first answer within four seconds before deliberating; final accuracy was 24.5%, but 67.2% of the people who finally answered correctly had already given the correct answer at that four-second first response. So for many solvers the right answer is an intuition too, not a rescue by slow reasoning. That is new context for anyone who treats the CRT as a clean divider between fast and slow thinkers, and our blog guide to why smart people make bad decisions touches on the same gap between being able to reason and doing it.
Can AI solve the Cognitive Reflection Test?
Chatbots have now overtaken people on it. Hagendorff, Fabi and Kosinski (2023), writing in Nature Computational Science, ran 150 new CRT-style tasks: ChatGPT-4 got 96% right, ChatGPT-3.5 got 59%, an earlier GPT-3 model got 5%, and the human comparison figure they used was 38%. A model that answers the standard bat-and-ball question correctly may simply have read the answer, so the new tasks matter more than the famous three. Our guides to the PNAS reasoning benchmark and to why AI benchmark scores disagree cover how fragile such comparisons are. For a different human-versus-machine test, see our guide to whether AI has passed the Turing test.
Where the Cognitive Reflection Test is used, and what it cannot tell you
Behavioural economists use the CRT to explain differences in patience, risk taking and susceptibility to biases. A 2018 trading-experiment study in the Journal of Finance, covered in our guide to what quant trading firms look for, found cognitive reflection was the strongest single predictor of avoiding the behavioural biases that erode a trader's returns, ahead of fluid intelligence on its own. As an assessment of a person it is thin: three items, four possible scores, widely leaked questions and no norm table. It measures a habit of checking, plus some number skill, plus a little memory. If you want a result on the standard IQ scale with a percentile, use a full test.
Where would your own score land?
Want a score on the standard IQ scale, with its percentile and a full report? Take the IQ Metrics test.
Find your IQ score now! →To see where a result of your own lands, take the IQ Metrics test and read it against the bell curve. For the general question of what a score tells you, our guide to how rare a high IQ score is puts the numbers in proportion.
Common questions
What is the answer to the bat and ball problem?
The ball costs 5 cents and the bat costs $1.05. At 10 cents the bat would be $1.10 and the pair would cost $1.20, not $1.10.
What is a good Cognitive Reflection Test score?
There is no pass mark. In Frederick's pooled sample of 3,428 respondents the mean was 1.24 out of 3, so 2 or 3 correct is in the top 40% and 3 is in the top 17% of that sample. MIT students averaged 2.18.
Is the Cognitive Reflection Test an IQ test?
No. It has three items and no IQ conversion table. A 2022 meta-analysis found it correlates r = .38 with cognitive intelligence, .61 after corrections, and .79 with numeracy.
Does seeing the CRT before change your score?
Yes. In one study of 142 people, those who had seen an item scored 2.36 against 1.48 for those who had not. Later work found the test still predicts reasoning measures after repeat exposure.
Sources for this story
- Cognitive Reflection and Decision Making (Frederick, 2005), Journal of Economic Perspectives 19(4):25-42 — American Economic Association
- Cognitive reflection, cognitive intelligence, and cognitive abilities: a meta-analysis (Otero, Salgado and Moscoso, 2022), Intelligence 90:101614 — Elsevier
- The Cognitive Reflection Test as a predictor of performance on heuristics-and-biases tasks (Toplak, West and Stanovich, 2011) — Memory & Cognition
- The smart System 1: evidence for the intuitive nature of correct responding on the bat-and-ball problem (Bago and De Neys, 2019) — Thinking & Reasoning
- Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in ChatGPT (Hagendorff, Fabi and Kosinski, 2023) — Nature Computational Science
- CRT familiarity studies: Haigh (2016), Stieger and Reips (2016), Bialek and Pennycook (2018) — Advances in Cognitive Psychology; PeerJ; Behavior Research Methods
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
Related stories
All news →
What IQ Do Quant Trading Firms Look For? Not a Number
Jane Street opens its process with a timed mental-math test. Renaissance Technologies built the most successful hedge fund in history hiring physicists over finance graduates. Neither one runs an IQ test.

Is IQ Just Pattern Recognition? What the Test Items Show
Pattern-finding is one ingredient of an IQ score, not the whole recipe. On the WAIS-5, two of the seven Full Scale subtests are built around it, and the factor data show why that is not enough.

Cattell Culture Fair Intelligence Test Explained
Raymond Cattell built the CFIT in 1949 to measure fluid reasoning without leaning on language or schooling. It is still used in cross-cultural research today, its four nonverbal subtests are unlike anything on a Wechsler scale, and its scoring uses a standard deviation of 24, not the usual 15.
Read the research.
Then find your own number.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →
