Stroop Test Explained: What It Measures and the IQ Link
In Stroop's 1935 study, naming the ink colour of 100 conflicting colour words took 47 seconds longer than naming coloured squares. The effect is among the most reliable in psychology, yet one person's score is not, and it tracks IQ only weakly.

What you need to know
- In John Ridley Stroop's 1935 experiments, 100 students took 110.3 seconds on average to name the ink colours of 100 conflicting colour words, against 63.3 seconds for 100 plain colour squares, a rise of 74%. Reading the same words printed in conflicting colours added only 2.3 seconds (5.6%) in a separate group of 70 undergraduates, and Stroop judged that gap not reliable.
- The Stroop effect appears in almost every group, but a single person's score is much noisier. In a 2018 test-retest study of the computer Stroop, the reaction-time cost had a reliability of .60 and .66 in two samples (47 and 56 adults, three weeks apart), with a standard error of measurement of 21 and 24 ms.
- The Stroop test is not an IQ test and no Wechsler subtest is a Stroop task. In a latent-variable study of young adults, updating working memory correlated highly with intelligence while inhibiting prepotent responses was small and not significant, and in 201 undergraduates fluid intelligence predicted a composite of three interference tasks.
- Stroop scores differ between groups with ADHD and without, but the size depends on how the test is scored: a 2007 meta-analysis found an average effect of 0.24 on difference scores and 1.11 in time-per-item studies. A group difference of that kind cannot diagnose one person.
The Stroop test asks you to name the colour of the ink a word is printed in while ignoring what the word says. When the word RED is printed in blue ink, people are slower and make more mistakes than when it is printed in red. That extra time or error rate is the Stroop effect, and it is one of the most replicated findings in experimental psychology. It is also a test of attention and speed rather than of intelligence, and its popularity online has made a lot of people ask what a good Stroop score is and whether it says anything about IQ.
How does the Stroop test work?
A Stroop task compares conditions. In the congruent condition the word and the ink agree (RED in red). In the incongruent condition they clash (RED in blue). Many versions add a neutral condition, such as a row of coloured squares or a word that is not a colour, as a baseline. Your Stroop cost is the incongruent time minus the congruent or neutral time. It is a difference score, a number made by subtracting one measurement from another, and that matters later in this article.
Reading is so practised that the word is processed whether you want it to be or not, which is why you cannot simply switch it off. It is a useful picture of how an automatic habit competes with an instruction. The same logic runs through other timed tasks on this site, such as the reaction time test and the digit span test, which measure speed and short-term memory instead of conflict.
What did Stroop's 1935 experiments find?
John Ridley Stroop published two experiments in the Journal of Experimental Psychology in 1935. Both used printed sheets of 100 items read aloud. The first group was 70 college undergraduates (14 men and 56 women), and the second was 100 students: 88 undergraduates and 12 graduate students, who were all women. The numbers from the paper are:
- Reading 100 colour names printed in black: the baseline for reading
- Reading 100 colour names printed in clashing ink colours: 2.3 seconds longer, or 5.6%, a difference Stroop called not reliable
- Naming 100 coloured squares: a mean of 63.3 seconds
- Naming the ink colour of 100 clashing colour words: a mean of 110.3 seconds, a rise of 47.0 seconds or about 74%
The asymmetry is the finding. A clashing ink colour barely slows reading, but a clashing word slows colour naming by almost three-quarters, because reading is the stronger habit. Stroop also reported that the standard deviation rose from 10.8 to 18.8 seconds, in step with the mean, so the coefficient of variability stayed at .171 in both tasks. In plain terms, the spread of times grew in proportion to the average, so a raw difference in seconds is larger for slower people, which is one reason later researchers argued for ratio scores.
What does the Stroop test measure?
It is usually described as a measure of interference control or response inhibition: holding on to a goal (name the colour) while suppressing a stronger habit (read the word). A 2018 paper by Hedge, Powell and Sumner groups the Stroop with the flanker, go/no-go and stop-signal tasks as tasks considered to be measures of impulsivity, response inhibition or executive functioning.
That label is itself contested. In a 2020 study of 201 college undergraduates, Paap and colleagues gave four non-verbal interference tasks, including a spatial version of the Stroop, and found that the scores did not correlate with self-reported self-control or impulsivity. They concluded that those tasks should not be read as measures of domain-general inhibitory control. A spatial Stroop is not the classic colour-word test, so this does not settle the colour-word version, but it shows how much weight a single interference score can carry.
Is a Stroop score reliable?
For groups, yes. Hedge and colleagues point out that the Many Labs 3 replication project found the Stroop effect in 100% of its attempts. For one person, much less so. They retested adults three weeks apart on seven classic tasks and found reliabilities from 0 to .82. The Stroop reaction-time cost came out at:
- Reliability (ICC) of .60 in Study 1 and .66 in Study 2, each with a wide range of uncertainty (.31 to .78 and .26 to .83)
- Standard error of measurement of 21 ms in Study 1 and 24 ms in Study 2
- Samples of 47 and 56 adults after exclusions, using four colours (red, blue, green, yellow) and key presses rather than spoken answers
The authors call this the reliability paradox. A task becomes a robust experimental effect when nearly everyone shows it to a similar degree, and the same low variation between people is what makes it hard to tell one person from another. Subtracting one time from another also adds noise, since a difference score has less usable variation than either of its parts. The Stroop cost was among the better scores in the study, since only the Stroop cost and go/no-go commission errors passed their .6 mark in both samples, yet a standard error of 21 to 24 ms means one session's result typically sits that far from the person's average. For comparison, a full-scale IQ score carries a standard error of about 2 to 3 points on the 15-point scale.
The Stroop effect is reliable for a group, but one person's Stroop score is a noisy number and should not be read as a trait.
Does the Stroop test relate to IQ?
Weakly, and the direction depends on which executive skill you ask about. Friedman and colleagues (2006) tested young adults on three executive functions and on fluid, crystallized and Wechsler IQ. Updating working memory correlated highly with all three intelligence measures, but inhibiting prepotent responses and shifting mental sets did not, and in models that controlled for the overlap between them, the inhibiting and shifting links to intelligence were small and not significant. In that study, the executive skill tied to IQ was updating working memory, not suppressing a habit.
Paap's 2020 data show a link for a related family of tasks: fluid intelligence, measured with a 12-item matrix test, was a significant predictor of the composite interference score, as was sex, while bilingualism, music training, video gaming, mindfulness and exercise were not. We did not find a large-sample study that correlates the classic colour-word Stroop with a full-scale IQ score, so any single figure quoted online should be treated with caution.
There is also a structural reason a Stroop result is not an IQ. The Wechsler scales build the full-scale score from subtests such as Vocabulary, Similarities, Block Design, Matrix Reasoning and Digit Span, and none of them is a Stroop task; see our guides to WAIS-IV and WAIS-5 index scores and to what the test items actually ask. Speed does matter a little: choice reaction time correlated about -.44 to -.53 with a timed intelligence test in 2,196 people, as we report in our reaction-time article. A Stroop trial is a choice reaction with conflict added, so some of what it measures is that same raw speed, and a pure speed measure still explains well under half the variance in IQ.
Can the Stroop test diagnose ADHD?
No. Lansbergen, Kenemans and van Engeland (2007) pooled the studies and found that the answer depends on how the test is scored. The mean effect size for ADHD relative to controls was 0.24 across all studies using difference scores, but 1.11 in studies that scored time per item. Using ratio scores, which are less sensitive to the outcome variable used, 19 studies showed more interference in the ADHD groups, and the authors conclude that interference control is consistently compromised in ADHD, a finding about groups.
Our arithmetic shows why that cannot make a diagnosis. An effect of 0.24 standard deviations puts the average person with ADHD at about the 59th percentile of the comparison group, and 1.11 at about the 87th. Those distributions overlap heavily, so many people with ADHD score typically and many people without it score slowly. If attention is a concern, an assessment by a clinician is the route, and our guides to ADHD and IQ and ADHD, autism, dyslexia and IQ test scores explain how cognitive testing fits in.
How should you read an online Stroop result?
There is no universal norm table for online Stroop tests, because the numbers depend on the set-up. Stroop's sheets were read aloud and timed for 100 items; Hedge's version used four colours and keyboard presses; other sites use two or three colours, mouse clicks or touch screens. Compare your own congruent and incongruent times on the same page and the same device, not against a figure from another source. A few habits make the result less noisy:
- Take it a few times on one device and average the incongruent minus congruent cost, since one run is noisy
- Watch accuracy as well as speed: a fast run with several wrong answers is not a better result
- Test at the same time of day and avoid doing it with a notification pinging
- Treat a change of a few tens of milliseconds as inside the noise, based on the 21 to 24 ms standard error above
- Do not convert it to an IQ: no published conversion exists, and a Stroop cost is not on the IQ scale
Stroop-type games sometimes appear in brain-training apps, and the evidence on whether brain training transfers to other skills is thin. The same lesson applies to practice on the test itself. You get better at the version you practise, which says little about how well you handle conflict elsewhere. If you want to see how a result of your own compares with the population, the IQ percentile calculator and the IQ bell curve put scores on a common scale.
Where would your own score land?
Want a result on the standard IQ scale instead of a millisecond cost? Take the IQ Metrics test and read your score against the bell curve.
Find your IQ score now! →A Stroop task is best treated as a clean demonstration of attention and habit, not as a measure of intelligence. For a puzzle-style read on reasoning, try the IQ Metrics test, and see how the related Cognitive Reflection Test asks you to override a gut answer, a close cousin of the Stroop conflict, and how dual n-back training targets the working-memory skill that does track IQ.
Common questions
What does the Stroop test measure?
It measures how much a clashing, automatic response (reading a word) slows you when you must do something else (name the ink colour). It is usually described as interference control or response inhibition, and it is also a speed task.
What is a good Stroop test score?
There is no universal norm, because results depend on the set-up. In Stroop's 1935 sheets, naming 100 clashing words took 110.3 seconds against 63.3 for colour squares. Compare congruent and incongruent times on the same device, and average several runs.
Is the Stroop test an IQ test?
No. None of the Wechsler subtests is a Stroop task, and a latent-variable study found inhibiting prepotent responses was not significantly related to intelligence once updating working memory was accounted for. Choice reaction time does correlate with IQ, but weakly.
Can the Stroop test diagnose ADHD?
No. A 2007 meta-analysis found larger interference in ADHD groups, with an effect of 0.24 on difference scores and 1.11 in time-per-item studies, but that is a group average and the score ranges overlap heavily, so it cannot diagnose one person.
Sources for this story
- Studies of interference in serial verbal reactions. Journal of Experimental Psychology, 18(6), 643-662 (1935), via Classics in the History of Psychology — John Ridley Stroop
- The reliability paradox: why robust cognitive tasks do not produce reliable individual differences. Behavior Research Methods, 50(3), 1166-1186 (2018) — Hedge, Powell and Sumner
- Not all executive functions are related to intelligence. Psychological Science, 17(2), 172-179 (2006) — Friedman, Miyake, Corley, Young, DeFries and Hewitt
- Interference scores have inadequate concurrent and convergent validity. Cognitive Research: Principles and Implications, 5, 7 (2020) — Paap, Anders-Jefferson, Zimiga, Mason and Mikulinsky
- Stroop interference and attention-deficit/hyperactivity disorder: a review and meta-analysis. Neuropsychology, 21(2), 251-262 (2007) — Lansbergen, Kenemans and van Engeland
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
Related stories
All news →
Average Reaction Time by Age, and What It Says About IQ
A typical simple reaction time is roughly 200 to 270 milliseconds, but the device and the task move it more than age does. In UK samples it correlates about -.3 with a timed intelligence test.
Working Memory and Reasoning Are Closely Related. They Are Not the Same Thing
The correlation between them is one of the sturdier findings in cognitive psychology. How strong it is, and what it implies about training either one, is where the agreement stops.

Cognitive Reflection Test: What the Bat-and-Ball Score Says About IQ
Across 3,428 people the three-question Cognitive Reflection Test averaged 1.24 correct, and only 17% got all three. Its link to IQ is moderate, not one-to-one, and the questions are now so well known that familiarity shifts scores.
Read the research.
Then find your own number.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →
