IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

Multitasking and IQ

Research & Evidence

Does Multitasking Lower Your IQ? What the Research Shows

No study has measured a person’s IQ before and after a bout of multitasking. What researchers at Stanford measured instead is narrower and more surprising: heavy multitaskers, tested against the exact skill they practice constantly, turned out to be worse at it, not better.

Bar chart comparing task-switching response times between heavy and light media multitaskers: heavy multitaskers 426 milliseconds slower on switch trials and 259 milliseconds slower on nonswitch trials, Stanford study 2009

No study has measured a person’s IQ before and after a bout of multitasking and found it dropped, and none has followed people over years to see whether heavy multitasking causes a lasting change in general intelligence. What researchers have measured is narrower and, in its own way, more interesting: chronic heavy multitaskers perform measurably worse on the exact cognitive skills an IQ test leans on most — filtering out irrelevant information, holding items in working memory, and switching cleanly between tasks. The researchers who first tested this in the lab expected the opposite result, which is a large part of why the finding is worth taking seriously.

The study that expected the opposite result

A 2009 Stanford study published in the Proceedings of the National Academy of Sciences set out to answer a specific question: do people who chronically juggle several media streams at once develop better cognitive control as a result of all that practice? The researchers built a media multitasking index from a questionnaire covering twelve forms of media, then compared "heavy" multitaskers (one or more standard deviations above the average number of simultaneous streams) against "light" multitaskers (one or more below) on a battery of established cognitive-control tests. The working hypothesis, reasonable on its face, was that heavy multitaskers should show an advantage at managing multiple streams of information, since that is exactly what they do all day.

What they actually found

In a filter task requiring participants to track two target shapes while ignoring a variable number of distractor shapes, heavy multitaskers’ accuracy fell steadily as distractors were added; light multitaskers’ accuracy did not move at all, meaning they filtered out the irrelevant shapes completely while heavy multitaskers let more and more of them in. A related test asked participants to hold a letter cue in mind and respond only when a specific follow-up letter appeared, sometimes with an unrelated distractor letter shown in between. With no distractor present, the two groups performed identically. Adding the distractor changed that: heavy multitaskers slowed down noticeably while accuracy stayed the same for both groups, meaning the extra letter was genuinely pulling their attention rather than confusing either group about the rules of the task. A separate test of raw impulse control, a stop-signal task requiring participants to withhold an already-triggered response, found no difference between the groups at all — heavy multitaskers were not simply more impulsive across the board, only more susceptible to letting irrelevant information in. A third test, a memory-updating task requiring participants to ignore letters that had appeared earlier but were no longer relevant, found the same pattern: heavy multitaskers’ false-alarm rate — mistaking an old, irrelevant item for a current target — climbed faster as the task went on. Across three independently designed tests, the direction of the effect was consistent: heavy multitaskers let more irrelevant information into working memory, not less, while showing no general deficit in impulse control on its own.

Bar chart comparing task-switching response times between heavy and light media multitaskers: heavy multitaskers 426 milliseconds slower on switch trials and 259 milliseconds slower on nonswitch trials, Stanford study 2009
Bar chart comparing task-switching response times between heavy and light media multitaskers: heavy multitaskers 426 milliseconds slower on switch trials and 259 milliseconds slower on nonswitch trials, Stanford study 2009

The task-switching result specifically

The most direct test compared how much slower each group got when a trial required switching from one type of task to a different one, versus repeating the same task type. If heavy multitasking practice built genuine switching skill, heavy multitaskers should have shown a smaller penalty. They showed a larger one: 167 milliseconds greater than light multitaskers’, a statistically solid difference. The detail that matters most is that heavy multitaskers were slower even on trials that did not require switching at all — 259 milliseconds slower on repeat trials, on top of 426 milliseconds slower on switch trials. That rules out an explanation limited to "switching itself is the problem"; something about sustained, single-task focus was affected too, not only the transition between tasks.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

What this does not show

The study’s own authors were explicit that the direction of cause and effect is unresolved, and this article carries that caveat forward rather than smoothing it over: chronic multitasking could cause weaker filtering and switching ability, or people who are already less able to filter distraction could simply be drawn toward multitasking behavior more than people who filter well. A follow-up summary of a decade of subsequent research, published by Stanford in 2018, described a consistent memory-performance gap between heavy and light multitaskers across many later studies — and reported the same causal-direction question still unresolved a decade later. The sample in the original study was also a group of university students tested once, not a general population followed over time, which limits how far any single number here should be generalized. The researchers themselves floated a specific version of the reverse-causation possibility worth stating directly: someone who already struggles to sustain focus on one thing might find single-tasking uncomfortable and gravitate toward switching between several activities precisely because it suits a mind that already has trouble filtering, rather than multitasking behavior creating that difficulty from scratch. Either story is consistent with the same data, and only a study that follows the same people over years, watching multitasking habits and cognitive-control ability change together or apart, could tell them apart. That study has not been done yet.

What the switching cost looks like outside the lab

A separate line of research, distinct from the Stanford trait-based comparison above, has measured the practical cost of interrupting one task to handle another. The American Psychological Association, summarizing task-switching research, estimates that switching between tasks can consume up to 40 percent of someone’s otherwise-productive time. A 2009 study on what researchers call "attention residue" found that a portion of attention stays with an unfinished task even after switching away from it, measurably impairing performance on whatever comes next. Neither of these findings is about IQ or general intelligence either — they describe a real, practical cost to interrupting focused work, on top of, and separate from, the cognitive-control differences described above. A widely cited 2005 field study of office workers found that more than half of observed work tasks were interrupted before completion, with an average delay of roughly 25 minutes before the original task resumed. That number describes a single workplace observational study, not a laboratory measure of cognitive ability, and it varies enormously by job and by how the interruption is defined; it is included here as context for how large the practical cost of switching can look outside a controlled experiment, not as a second data point for the cognitive-control findings above.

So, does multitasking lower your IQ

Not in any sense a test can measure directly, and nobody has run the study that would settle it either way. What the evidence does support is narrower and arguably more useful: treating multitasking as a skill that improves with practice, the way learning an instrument or a language does, is not what the data shows. The Stanford research points toward the opposite pattern — constant practice at dividing attention correlating with getting worse, not better, at the specific mental skills, filtering distraction and holding focus, that an IQ test also draws on. Whether that is cause, effect or a bit of both for any one person is a question this research has not answered yet.

Where this fits next to the site’s other coverage

This is a different mechanism from this site’s piece on AI and cognitive offloading, which covers what happens when someone delegates a single thinking task to a tool rather than doing it themselves; multitasking is about dividing attention across several tasks at once, a distinct question with its own separate research base. It sits more naturally next to processing speed and IQ and memory and IQ as another way a specific, trainable-feeling mental skill turns out to have a more complicated relationship to general cognitive performance than intuition suggests. The same underlying theme — a limited cognitive resource, spent well or poorly — shows up from a completely different angle in this site’s piece on sleep and IQ, which is about losing hours of consolidation time rather than splitting attention during the day.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged attention, cognitive control, cognitive science, digital age cognition, distraction, executive function, intelligence research, IQ Science, media multitasking, multitasking, productivity, Stanford study, task switching, Working Memory

Prenatal Smoking and IQ

Research & Evidence

Prenatal Smoking and IQ: A Story More Complicated Than the Headlines

Several studies link smoking during pregnancy to a lower child IQ score. Then a Danish cohort of 1,782 mothers adjusted for the one variable most of the earlier research had not: the mother’s own IQ. What was left afterward is a genuinely more complicated story than the headline number suggests.

Bar chart comparing the reported IQ deficit from heavy prenatal smoking before and after adjusting for confounders: 4 points lower unadjusted, not statistically significant after adjusting for maternal IQ and other confounders, Danish National Birth Cohort 2012

Probably a real effect, but a smaller and more contested one than the plainest version of this claim suggests — and a meaningful share of what early research measured as a smoking effect looks, in the best-controlled study available, like something else entirely. Nicotine is a well-established neuroteratogen with a real biological pathway into a developing brain. It is also true that mothers who smoke during pregnancy differ, on average, from mothers who do not in ways that independently affect a child’s IQ score — and one of the more careful studies on this question found that once those differences are accounted for, the specific IQ-point effect mostly disappears.

Both of those things are true at once, and treating this as a simple "smoking lowers IQ by N points" claim, or dismissing it as pure confounding, both oversell the certainty the evidence actually supports.

The biological case, independent of any single study

Nicotine and its metabolite cotinine cross the placenta freely, reaching the fetus at concentrations equal to or higher than those in the mother’s own bloodstream. Both are established to alter the developing brain’s acetylcholine, serotonin and catecholamine neurotransmitter systems, and smoking also produces vasoconstriction and reduced oxygen delivery to the fetus through a separate pathway. None of that depends on any particular cohort study’s numbers holding up; it is why researchers expected to find a cognitive effect in the first place, and it is also why prenatal smoking’s other, non-cognitive harms — preterm delivery, fetal growth restriction, congenital malformation, stillbirth and Sudden Infant Death Syndrome — are not in dispute here at all. This article is about the narrower and genuinely contested IQ-point question only.

It is worth being precise about what "no known safe amount" means when public-health agencies say it, since the same phrase shows up in this site’s companion piece on prenatal alcohol exposure and is easy to misread. It is a statement that no safe threshold has been established by research, not a claim that every level of exposure produces a measurable cognitive effect. Heavy, sustained smoking carries well-documented risk across multiple outcomes; the size of any effect from lighter or more occasional exposure specifically on IQ is one of the genuinely less settled questions this article covers, and treating the two as identical overstates what either literature actually shows.

What the early cohort studies found

A cohort of more than 1,800 Estonian schoolchildren, drawn from 45 schools across all fifteen of the country’s counties, found a 3.3-point IQ deficit associated with prenatal smoking exposure, alongside separate effects from birth weight and maternal education. A 2012 study of the Danish National Birth Cohort, 1,782 mother-child pairs tested with a standard preschool IQ scale at age 5, found an unadjusted 4-point drop in Full-Scale IQ associated with smoking 10 or more cigarettes a day during pregnancy, compared with not smoking at all — a statistically solid, easily quotable number, and the kind of figure that tends to be the one repeated in summaries of this research. Both studies controlled for at least some confounding factors, which is standard practice in this field and is exactly why their headline numbers read as credible on their own. The open question was never whether these researchers adjusted for anything — it is whether they adjusted for the single confounder that turns out to matter most.

What happened when researchers controlled for the mother’s own IQ

That same Danish study did something a large share of earlier research had not: its full statistical model adjusted not just for parental education but for maternal IQ directly, along with maternal alcohol use, the child’s sex and age, paternal smoking, maternal age and body mass index, family environment, breastfeeding and sensory impairment. Once all of that was accounted for, the significant effect of prenatal smoking on Full-Scale IQ disappeared. The authors’ own reading is specific and worth quoting in substance rather than softening: earlier studies that did not control for maternal IQ directly may carry substantial residual confounding, because smoking during pregnancy correlates with a mother’s own cognitive score for reasons that have nothing to do with nicotine’s effect on a fetus.

Bar chart comparing the reported IQ deficit from heavy prenatal smoking before and after adjusting for confounders: 4 points lower unadjusted, not statistically significant after adjusting for maternal IQ and other confounders, Danish National Birth Cohort 2012
Bar chart comparing the reported IQ deficit from heavy prenatal smoking before and after adjusting for confounders: 4 points lower unadjusted, not statistically significant after adjusting for maternal IQ and other confounders, Danish National Birth Cohort 2012
Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

The complication that keeps this from being a clean confounding story

If confounding fully explained the pattern, a well-designed genetic-interaction study should not find a smoking-specific effect tied to one particular gene variant. One does. Children carrying the Met allele of the BDNF gene showed a Full-Scale IQ deficit of more than 8 points specifically when their mother smoked during pregnancy, a pattern that held across multiple childhood ages in the same cohort. A socioeconomic or maternal-IQ confound would be expected to affect children regardless of which BDNF variant they inherited; a gene-by-exposure interaction this specific is harder to wave away as pure confounding, and it is the strongest piece of evidence that a real, biological, dose-sensitive effect exists underneath the more contested average finding. Taken together, the two results are less contradictory than they first look: the average effect across an entire cohort may genuinely wash out once maternal IQ is accounted for, while a real, biologically specific effect still concentrates inside a genetically identifiable subset of the same cohort. A population-level null result and a subgroup-level real effect can both be true about the same dataset at once, a nuance no single quoted number ever carries on its own.

Secondhand exposure and dose

The research on secondhand smoke exposure during pregnancy follows a similar dose-response shape to the direct-smoking literature: more frequent secondhand exposure is associated with worse outcomes than occasional exposure, in the same general direction as, though generally smaller than, active maternal smoking. The practical implication tracks the direct finding above: dose and frequency appear to matter more than exposure as a simple binary, which is also consistent with the biological picture — a fetus absorbing a smaller, less sustained dose of nicotine and cotinine would be expected, on the mechanism described earlier, to show a smaller effect than one absorbing a mother’s own direct, daily exposure. It is also one more reason a single number attached to "smoking in pregnancy" as a yes-or-no category was always going to be an oversimplification of what is really a graded, dose-dependent exposure, closer in shape to a dial than to a simple binary light switch.

What this means, and does not mean

The honest summary sits between the two headline-ready versions of this story. Nicotine’s biological pathway into a developing brain is real and well established. Some of the specific IQ-point numbers attached to that pathway, in studies that did not control for the mother’s own cognitive ability, are very likely inflated by confounding — the same shape of caution this site applies to the Adverse Childhood Experiences score in the piece on toxic stress and IQ, where a real population-level signal turned out to predict very little about any individual person. And a specific gene-by-environment finding suggests the confounding explanation is not the whole story here either. Readers looking for a single settled number will not find one in this article, and that absence is itself the honest finding: unlike some of the other exposures covered on this site, the research here has not converged on one, and a responsible summary says so rather than repeating whichever figure sounds most citable. None of this touches prenatal smoking’s separately well-established harms outside cognition, which are not in question and are not softened by anything above. This is one of three ways this site has now covered a substance or exposure reaching a fetus directly: alongside prenatal alcohol exposure and, through the lungs rather than the placenta, fine-particle air pollution. For the much broader question of how much of any IQ score is inherited in the first place, including from a mother’s own cognitive ability, see is IQ genetic.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged birth cohort study, child iq, cognitive development, confounding variables, early childhood development, fetal development, gene environment interaction, intelligence research, IQ Science, maternal smoking, neuroteratogen, nicotine exposure, pregnancy health, prenatal smoking, tobacco

Air Pollution and IQ

Research & Evidence

Air Pollution and IQ: What the Research Actually Shows

A 2024 meta-analysis pooling six studies and 4,860 children found that every rise in fine-particle air pollution came with a small, statistically significant drop in IQ scores. The same pattern shows up in aging adults living near traffic, and in a household exposure route this conversation usually leaves out entirely.

Bar chart of IQ point change per 1 microgram per cubic meter increase in PM2.5 air pollution: full-scale IQ down 0.27, performance IQ down 0.39, verbal IQ down 0.24, from a 2024 meta-analysis of 4,860 children

Yes, and the evidence is more consistent than most people expect for a topic this new. A 2024 meta-analysis pooling six studies and 4,860 children across three continents found that every increase in fine-particle air pollution was associated with a small, statistically significant drop in IQ scores. The important word there is small: this is a population-level pattern, visible when averaging across thousands of children, not a prediction about what pollution did to any one child. Six included studies is also a genuinely thin evidence base for a meta-analysis, worth saying plainly rather than treating the finding as settled.

Fine particulate matter — PM2.5, particles smaller than 2.5 micrometers, small enough to lodge deep in the lungs and, a growing body of research suggests, to reach the bloodstream — sits alongside lead as one of the more actively studied environmental exposures in child cognitive development. The two are not the same exposure and do not share a source, but they raise a similar question: what does an involuntary, population-wide exposure do to a developing brain.

What the largest analysis to date found

Published in the journal Environmental Health in 2024, the meta-analysis screened 1,107 publications against PRISMA guidelines across seven databases and found six studies that met its inclusion criteria: 4,860 children across North America, Europe and Asia, exposed to a mean PM2.5 concentration of 30.4 ± 24.4 micrograms per cubic meter, tested at an average age of 8.9. Using a random-effects model, the researchers calculated that each 1 microgram-per-cubic-meter increase in PM2.5 was associated with a 0.27-point drop in Full-Scale IQ (p<0.001), a 0.39-point drop in Performance IQ (p=0.003), and a 0.24-point drop in Verbal IQ (p=0.021). Performance IQ, the nonverbal and perceptual-reasoning subtests, showed the largest and most consistent hit of the three — the studies included do not settle why, though a similar pattern of nonverbal and processing measures being the most exposure-sensitive shows up elsewhere in the environmental-neurotoxin literature.

Bar chart of IQ point change per 1 microgram per cubic meter increase in PM2.5 air pollution: full-scale IQ down 0.27, performance IQ down 0.39, verbal IQ down 0.24, from a 2024 meta-analysis of 4,860 children
Bar chart of IQ point change per 1 microgram per cubic meter increase in PM2.5 air pollution: full-scale IQ down 0.27, performance IQ down 0.39, verbal IQ down 0.24, from a 2024 meta-analysis of 4,860 children

How something this small could reach a developing brain

PM2.5 is defined by size, not by a single chemical identity — it is a mix of combustion byproducts, metals and organic compounds small enough to bypass the lungs’ normal filtering and settle deep in the smallest airways. From there, the leading proposed mechanisms are a systemic inflammatory response that reaches the brain through the bloodstream, and a separate, more direct route in which the smallest particles are thought to travel along the olfactory nerve into brain tissue itself, bypassing the blood-brain barrier entirely. Neither pathway is fully settled science, and researchers are still working out how much each contributes relative to the other. What is better established is the downstream signature: postmortem and imaging studies have found markers of neuroinflammation and, in some cases, combustion-derived particles themselves inside brain tissue, which is the kind of physical evidence that turns a statistical association into a plausible causal story rather than a coincidence.

Pollution at school, not just at home

A 2015 prospective study led by researchers at Barcelona’s Centre for Research in Environmental Epidemiology followed 2,715 children across 39 schools, testing their cognitive development four separate times over 12 months. The angle was deliberately different from most exposure research: many schools sit close to busy roads, and traffic pollution peaks during the exact hours children are inside them. Children attending higher-traffic-pollution schools showed measurably slower growth in working memory over that year than children at lower-pollution schools — a finding that points at a policy-addressable variable, where schools get built and how close to traffic, rather than only a home-level exposure families have less power to change.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

It does not stop at childhood

A separate and growing body of cohort research links long-term air pollution exposure to faster cognitive decline and higher dementia risk in older adults: a U.S. cohort study associating midlife air pollution and road proximity with incident dementia, a Swedish longitudinal study linking traffic-related pollution to dementia incidence, France’s long-running Three-City Study finding a similar association, and a Neurology-published analysis of cognitive-decline trajectories in older adults exposed to long-term air pollution. All four are observational cohort studies, not randomized trials, and carry the obvious open question any pollution-and-health cohort does: whether people living nearer to heavy traffic differ in other ways — income, housing quality, access to care — that could independently affect cognitive aging. Researchers adjust statistically for exactly these factors, which narrows but does not eliminate the concern. What makes the overall pattern harder to dismiss is that it shows up twice, independently, in childhood development data and in adult cognitive-aging data, using different cohorts, different countries and different outcome measures.

The exposure route this conversation usually skips

Nearly all popular coverage of air pollution and cognition focuses on outdoor, traffic-related PM2.5 in wealthy cities. The far larger exposure, by concentration, happens indoors. The World Health Organization’s 2021 guidelines halved the recommended annual PM2.5 limit from 10 to 5 micrograms per cubic meter; field measurements in poorly ventilated homes using wood or coal cookstoves have documented indoor PM2.5 concentrations around 5,000 micrograms per cubic meter — roughly a thousand times that guideline. Household air pollution is linked to an estimated 70 percent of all air-pollution-attributed deaths in children under 5, and young children are disproportionately exposed because they typically stay close to their mothers during cooking, absorbing concentrations far closer to a cook’s own exposure than to anything measured outdoors nearby. Read against that number, this is substantially a global-equity question, not only a question about traffic congestion in wealthy cities, and it is the exposure route least likely to come up in a conversation about air pollution and cognition even though it is, by concentration, by far the larger one.

What narrows the exposure

None of this is a reason for panic, and it is not a reason to treat every city as equally dangerous — exposure varies enormously by location, season and even time of day, and the effect sizes above are population averages, not individual predictions. A handful of concrete, evidence-backed steps do meaningfully reduce personal exposure, for a family or a school weighing what is actually worth changing:

  • Indoor air filtration (a HEPA filter measurably lowers indoor PM2.5 in homes near heavy traffic or wildfire smoke)
  • Checking a local air quality index before strenuous outdoor exercise, particularly near arterial roads
  • Ventilation improvements or cleaner-burning cookstoves, the single highest-leverage change in the household-exposure settings described above
  • Advocating for school siting and traffic-calming decisions, since the Barcelona finding above ties exposure directly to where a building sits

The honest bottom line

The size of the effect, point for point, is small enough that no single study proves anything about any individual child or adult. What makes environmental epidemiologists take it seriously anyway is that the same direction of harm turns up repeatedly: across a meta-analysis of six independent studies, in a dedicated school-exposure cohort, and separately in adult cognitive-aging research, using different populations and different measurement methods each time. Six studies feeding one meta-analysis is not a large evidence base, and a skeptical reader is right to want more before treating the exact numbers — 0.27, 0.39, 0.24 points per microgram — as fixed. The more defensible claim is the direction and the breadth of it: several independent lines of research, run by different teams on different continents using different children and different adults, keep landing on the same conclusion, which is a higher bar to clear by accident than any one of them would be alone.

That is a different, more airborne kind of early-life exposure than this site’s companion piece on prenatal smoking, which covers a chemically distinct toxin crossing the placenta directly rather than the lungs; both sit alongside the site’s piece on chronic early-life stress as reminders that a developing brain can be shaped by exposures that have nothing to do with inherited ability.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged air pollution, brain health, child iq, cognitive decline, cognitive development, dementia risk, environmental exposure, environmental health, household air pollution, intelligence research, IQ Science, neurotoxicity, particulate matter, PM2.5, traffic pollution

IQ Testing and Eugenics

Understanding IQ

Army Alpha and Beta: How IQ Testing Got Tangled Up With Eugenics

Between 1917 and 1919, the United States Army tested roughly 1.75 million recruits in the largest intelligence-testing program the world had yet seen. Within a decade, one psychologist had turned that data into a case for restricting immigration by race. Thirteen years after that, the same psychologist published a retraction few people have ever heard of.

Timeline of IQ testing and eugenics: 1917 Army Alpha and Beta tests administered, 1923 Brigham publishes A Study of American Intelligence, 1924 Immigration Act restricts immigration by national origin, 1930 Brigham retracts his own conclusions, 1981 Gould publishes The Mismeasure of Man

Closely, and unhappily, for a specific stretch of the early 20th century — though less directly, and less simply, than the popular version of this story usually has it. Between 1917 and 1919, the U.S. Army tested roughly 1.75 million recruits using two new instruments, Army Alpha and Army Beta, in what was at the time the largest intelligence-testing program ever conducted. Within a few years, a young psychologist named Carl Brigham had turned that wartime data into a book arguing for a fixed racial hierarchy of innate intelligence. Within a decade after that, Brigham published a retraction of his own conclusions that almost nobody who cites his 1923 book seems to know about.

That is the responsible, checkable version of a history that gets told two ways: either as a footnote nobody looks at closely, or as a much simpler morality tale than the actual documentary record supports. Both the original claims and the standard modern correction to them turn out to have real, specific problems worth naming.

The largest intelligence-testing program in history, to that point

The Army program was directed by the psychologist Robert Yerkes, with a committee of the era’s leading testers including Henry Goddard and Lewis Terman, the psychologist who would go on to adapt the Stanford-Binet scale. Army Alpha was a written, group-administered test of analogies, arithmetic and judgment for literate, English-speaking recruits; Army Beta was a nonverbal, pictorial equivalent built for recruits who were illiterate or did not speak English. The stated purpose was practical rather than ideological on its face: classify recruits for duty assignment, screen out roughly 8,000 men judged unfit for service, and help select officer candidates, ultimately factoring into the selection of something like two-thirds of the war’s 200,000 commissioned officers.

Ellis Island, and a study smaller than its reputation

A few years earlier, the psychologist Henry Goddard had sent researchers, and later went himself, to test immigrants arriving at Ellis Island, publishing the results in 1917. The study is still cited as if it were a broad survey of immigrant intelligence; it was nothing close to that. Government physicians had already pulled aside anyone visibly disabled before Goddard’s team arrived, and the researchers then set aside anyone who looked obviously capable as well, testing only a small, deliberately ambiguous remainder. Using a looser, group-adjusted scoring standard on that unrepresentative slice, Goddard reported that more than 40 percent of the Jewish immigrants he tested scored as "feebleminded." The study’s sampling problem, not just its conclusions, is the part worth remembering: a small, pre-filtered group testing poorly on an unfamiliar, culturally loaded exam in a language many of them did not speak was reported as if it described immigrants generally.

The book that turned Army data into a racial hierarchy

In 1923, Brigham published A Study of American Intelligence, using the wartime Army Alpha and Beta results to argue for a ranked hierarchy of racial and national groups, Nordic peoples at the top, and to warn that continued immigration was driving a decline in American intelligence. The book was taken seriously in exactly the circles working to restrict immigration at the time, and it gave that political project something it wanted badly: a veneer of hard, quantitative science behind an argument that was, underneath, about who counted as fit to become American.

Timeline of IQ testing and eugenics: 1917 Army Alpha and Beta tests administered, 1923 Brigham publishes A Study of American Intelligence, 1924 Immigration Act restricts immigration by national origin, 1930 Brigham retracts his own conclusions, 1981 Gould publishes The Mismeasure of Man
Timeline of IQ testing and eugenics: 1917 Army Alpha and Beta tests administered, 1923 Brigham publishes A Study of American Intelligence, 1924 Immigration Act restricts immigration by national origin, 1930 Brigham retracts his own conclusions, 1981 Gould publishes The Mismeasure of Man
Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

How much this actually shaped the Immigration Act of 1924

Here is where the popular version of this story runs ahead of the documentary record. A detailed 1983 review of the actual congressional hearings and floor debate behind the Immigration Act of 1924 found intelligence testing unmentioned in the law’s text, raised exactly once across more than 600 pages of floor debate, and criticized there without any rebuttal from supporters. None of the era’s prominent testers, not Yerkes, not Goddard, not Terman, testified before Congress or had their work entered into the official record. The reviewers’ own conclusion was that the test results were "surely not crucial" to the legislation and "most likely… immaterial" to it, a conclusion an independent 1982 review reached separately. None of that erases eugenic ideology from the story of the 1924 Act — a broader movement, and Brigham’s book specifically, were genuinely invoked in committee testimony supporting the restrictions. What the documentary record does not support is the stronger, more dramatic claim that Army intelligence-testing data was a direct or decisive cause of the law.

The retraction almost nobody remembers

In 1930, Brigham published a short paper in Psychological Review with a title that undersells what is inside it, "Intelligence Tests of Immigrant Groups." In it, he wrote that his own 1923 study — "one of the most pretentious of these comparative racial studies" — "was without foundation." His specific point was that comparing groups on a test given in English, when many test-takers had limited English fluency and limited schooling, measured language and educational access far more than it measured innate ability — exactly the kind of cultural and linguistic bias this site covers in a separate piece on test bias. It is worth being precise that the 1930 retraction focused on that methodological problem specifically; it is not a line-by-line repudiation of every claim in the 1923 book. One further, genuinely strange footnote: Brigham went on, in the years between the book and the retraction, to become the chief architect of the SAT, adapting it directly from the Army testing methodology he would later renounce for a very different purpose.

The standard modern critique, and a critique of the critique

The biologist Stephen Jay Gould’s 1981 book The Mismeasure of Man is the most widely read modern account of this history, arguing that the Army testing program was scientifically compromised, from testing conditions to culturally loaded content, and that the resulting data, filtered through Brigham, helped drive the 1924 restrictions. Two layers of scholarly pushback have followed, and they are worth keeping separate. The first, the 1983 congressional-record review described above, argues Gould overstated how directly the testing data shaped the actual legislation. The second is narrower and more recent: a 2019 paper in the Journal of Intelligence checked Gould’s specific claims about the Army Beta test itself against the original Army records and found real discrepancies — Gould described "vast numbers" of test-takers scoring zero, where the actual rate ran from under 1 to about 10 percent per subtest, and Gould cited a single unfavorable administration report while roughly a dozen more favorable ones existed in the same archive. That paper’s own conclusion is that Army Beta was, by the standards of 1917, a reasonably well-designed test. No published rebuttal of that specific finding exists yet. The pattern worth noticing is not that any one side was simply right: it is that a confident claim about what IQ data proved, made in 1923, was substantially corrected in 1930 by the person who made it, and a confident modern correction to that same history has itself since been checked, and partly corrected, against the primary record.

What actually holds up, a century later

Eugenic ideology, as a broader movement, was real, influential and genuinely tied to American immigration and sterilization policy in this period — that much is not in dispute among historians. What the specific documentary record does not support is treating IQ-testing data itself as the direct engine of that policy, rather than one piece of scientific-sounding cover recruited into an argument that was already being made on other grounds. The durable, checkable lesson is narrower than either the popular story or its detractors usually make it: a test built and interpreted carelessly, given in an unfamiliar language to people with unequal access to schooling, was read as measuring fixed, innate group differences it was never capable of measuring — and the field’s own literature, from Brigham’s 1930 paper onward, has been correcting that specific mistake in some form ever since, right up to modern work on what a heritable trait does and does not mean for group differences and to this site’s own piece on why national IQ-ranking claims do not support what people use them to argue, a modern descendant of exactly the same error. This is also not the only place on this site where a confident, well-published claim about what an IQ-adjacent measure proves did not survive later scrutiny — a very different, much more recent example plays out over years rather than decades in this site’s piece on the bilingual advantage. For the broader arc this one chapter sits inside, see this site’s general history of IQ testing.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged army alpha test, army beta test, carl brigham, Cultural Bias, eugenics, henry goddard, history of iq testing, immigration act of 1924, intelligence research, iq myths, IQ Science, robert yerkes, stephen jay gould, test validity

Bilingualism and IQ

Research & Evidence

The Bilingual Advantage: What the Executive-Function Research Actually Shows

For about a decade, a well-known finding held that bilingual children and adults show sharper executive function than monolinguals. Then researchers tried to replicate it across 15 separate measures and found almost nothing. A study of conference abstracts explained part of why the early evidence looked so much stronger than what came next.

Bar chart comparing the reported bilingual executive-function advantage in children before and after correcting for publication bias: 0.08 standard deviations before correction, falling to negative 0.04, statistically indistinguishable from zero, after correction, from Lowe et al. 2021

Probably not in the way a decade of headlines suggested, and the story of how researchers found that out is a clean, well-documented case of a plausible finding failing to survive closer scrutiny. The original claim, built on real studies by a respected research group, was that juggling two languages exercises a general-purpose control system in the brain, producing an advantage on tasks that require ignoring irrelevant information. A direct replication attempt across 15 separate measures of that exact skill found essentially nothing, and a large 2021 study of more than 23,000 children — the population the original claim was actually about — reduced the effect to zero once a specific statistical bias was corrected for.

None of that means the underlying research was done badly. It means this is one of the cleaner examples in psychology of how a real, appealing early finding can look much stronger in the published literature than it turns out to be once other labs go looking for it directly.

The original finding, and why it was plausible

The foundational work came from Ellen Bialystok and colleagues, first in preschoolers using card-sorting tasks that require switching rules, then extended to middle-aged and older adults using the Simon task, a reaction-time test where a stimulus’s position conflicts with the response it requires. Bilingual participants in these studies showed a smaller performance cost from that conflict than monolinguals, and the proposed explanation was straightforward: managing two active languages and suppressing the one not currently in use functions as a kind of constant low-level exercise for a domain-general executive-control system, one that should show up on any task drawing on the same underlying skill.

The replication that did not confirm it

In 2013, researchers set out to test that prediction directly and broadly, gathering roughly 15 separate indicators of executive control across four different task types — antisaccade, Simon, flanker and task-switching paradigms — in young adults. If the domain-general theory were right, a bilingual advantage should have turned up on most of them. It did not turn up as a significant effect on any of the 15 measures, and the single interaction that did reach significance pointed toward a disadvantage for bilinguals, not an advantage. The paper’s own title stated the conclusion plainly: there was no coherent evidence for a bilingual advantage in executive processing.

Bar chart comparing the reported bilingual executive-function advantage in children before and after correcting for publication bias: 0.08 standard deviations before correction, falling to negative 0.04, statistically indistinguishable from zero, after correction, from Lowe et al. 2021
Bar chart comparing the reported bilingual executive-function advantage in children before and after correcting for publication bias: 0.08 standard deviations before correction, falling to negative 0.04, statistically indistinguishable from zero, after correction, from Lowe et al. 2021

A second problem: which studies got published

A 2015 study in Psychological Science offered a specific explanation for why the published record looked so much more convincing than a direct replication did. The researchers tracked every conference abstract on bilingualism and executive control presented between 1999 and 2012, then checked which of those studies eventually appeared as full journal articles. Studies whose results fully supported a bilingual advantage were published far more often than studies with mixed or contrary results, even though the groups did not differ in sample size or statistical power — a textbook publication-bias pattern, where positive findings are simply more likely to make it into print than null ones.

What happens when you pool everything, published or not

A 2018 meta-analysis in Psychological Bulletin, deliberately built to include unpublished data alongside published studies, pooled 891 effect sizes from 152 studies of adults across six components of executive function. Before adjusting for publication bias, a very small advantage appeared in a few of those domains. After adjusting for it, the advantage disappeared in every domain tested, and a small disadvantage showed up for verbal fluency specifically, plausibly because bilinguals split their exposure to any one language across two rather than concentrating it in one.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

The same null result shows up in children now too

The strongest recent evidence extends the null finding to exactly the population the original claim was about. A 2021 study in Psychological Science analyzed 1,194 effect sizes from more than 23,000 children and adolescents: the raw advantage across that entire literature was 0.08 standard deviations, small but technically real. Once corrected for the same publication bias described above, that number fell to negative 0.04 — statistically indistinguishable from zero, including inside the specific "executive attention" subdomain the original theory predicted should show the clearest effect. A separate study of more than 4,500 9- and 10-year-olds in a large, nationally representative U.S. cohort reached the same conclusion using an entirely different dataset.

Where a real effect might still survive

None of this fully rules out an effect under narrower conditions than the original claim proposed. One line of newer research argues that the binary bilingual-versus-monolingual comparison used in essentially every study above may itself be part of the problem: a 2019 paper argues bilingual experience is better treated as a continuous spectrum — proficiency, age of acquisition, how immersive the environment is — than a single either-or label, so a null result on a binary comparison could still be masking a smaller, dosage-dependent effect concentrated at the far end of that spectrum. A related theory, the adaptive control hypothesis, proposes that what actually matters is not bilingual status at all but how someone uses their languages day to day — dense, moment-to-moment code-switching between two languages in the same conversation is a much heavier load on executive control than keeping two languages largely separated by context or setting, and most of the studies above did not distinguish between the two usage patterns. Neither idea has yet produced the kind of large, preregistered, bias-corrected confirmation that settled the null result above; both are the live, serious version of "there might still be something here," not a reflexive defense of the original claim.

Even the original researcher has revised the theory

In 2024, Bialystok herself published a theoretical revision in Trends in Cognitive Sciences, moving away from the original "skill transfer" account — the claim that bilingual experience directly boosts general executive function as a flat, always-on advantage — toward a narrower "adaptation" framework that predicts a benefit only under specific, high-demand conditions rather than as a general boost. That is not a minor footnote: it is the field’s founding researcher publicly updating the theory in response to exactly the pattern of evidence described above, which is roughly the scientific process working the way it is supposed to.

A different, more contested claim worth keeping separate

This site’s news desk has separately covered a different bilingualism claim, about delayed dementia symptoms in older adults, and it is worth being precise that these are not the same finding restated. Two retrospective clinic studies found Alzheimer’s symptoms appearing roughly four to five years later in bilingual patients already diagnosed with dementia. That is a different measure, symptom timing in people who already have the disease, from a healthy person’s risk of developing it at all — and when researchers restricted the analysis to the stronger prospective study designs rather than the retrospective ones, the pooled result showed no reduced risk of actually developing dementia. Retrospective clinic samples like the original two studies are also particularly prone to referral-pattern bias: which patients get diagnosed, and when, can differ systematically between bilingual and monolingual communities for reasons that have nothing to do with brain biology. If it holds up at all, this is a narrower and later-life claim than the childhood executive-function story above, not confirmation of it. A third, unrelated finding covers something else again: taking an IQ test in your second language lowers verbal-subtest scores for reasons that have nothing to do with underlying ability.

The pattern this fits

A skill developed through intensive practice failing to transfer into a broad cognitive advantage is not unique to bilingualism on this site — world-class memory athletes show the same disconnect between a trained skill and general reasoning ability, and the research on commercial brain-training programs, covered on this site’s news desk, turns up a strikingly similar gap between narrow, trained improvement and any broader intelligence gain. Skepticism toward "doing X trains your brain generally" claims is, again and again, the position the evidence keeps rewarding. It is a far gentler version of a pattern that shows up with much higher stakes in this site’s piece on IQ testing’s eugenics era: a confident, published claim about what a cognitive measure proves, later substantially walked back once other researchers looked more carefully.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged bilingualism, brain training, cognitive development, cognitive reserve, executive function, general intelligence, intelligence research, iq myths, IQ Science, meta-analysis, publication bias, replication crisis, second language, Working Memory