IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

Toxic Stress and IQ

Research & Evidence

Toxic Stress and IQ: What the Bucharest Orphanage Study Actually Found

Researchers in Romania did something that can never be done again for ethical reasons: they randomly assigned some institutionalized children to foster care and left others in the institution, then followed both groups for twenty years. The IQ gap that opened up has held at every check-in since. What has not held up as cleanly is the famous claim about exactly when the window closes.

Bar chart of average IQ by care condition at age 4.5 in the Bucharest Early Intervention Project: 109 for never-institutionalized children, 81 for children randomly assigned to foster care, and 73 for children who remained in institutional care

Yes, and the cleanest evidence available comes from the only randomized controlled trial ever conducted on the question — a study that could not be ethically repeated today and has followed the same children for two decades. Children living in Romanian institutions were randomly assigned either to remain in institutional care or to move into a newly created foster-care network. At the first major follow-up, the foster-care group averaged an IQ of 81, the group that remained institutionalized averaged 73, and a comparison group of children who had never been institutionalized averaged 109.

That eight-point gap between the two institutionalized groups — caused by nothing except which one got assigned to leave — is about as close to direct causal proof as developmental psychology gets. What is less solid, and worth being honest about, is the specific age-based claim that usually rides along with this study whenever it gets summarized.

The experiment nothing else in this field can match

The Bucharest Early Intervention Project enrolled 136 children, average age about 22 months, living in Romanian institutions in the early 2000s. Half were randomly assigned to continued "care as usual"; the other half were placed into foster families recruited and supported specifically for the study. A third group of Bucharest children who had never been institutionalized served as a community comparison. Because placement was randomized rather than chosen by caseworkers or families, any later difference between the foster-care and institutional groups can be attributed to the intervention itself rather than to whatever made some children more likely to be placed — the same logic, and the same rare strength, behind this site’s piece on the Perry Preschool experiment.

What the gap looked like, and how long it lasted

At 54 months, the foster-care group’s average IQ of 81 sat meaningfully above the institutional group’s 73, both well below the never-institutionalized comparison group’s 109. The gap did not fade the way some early-childhood effects do: at age 12, the same three groups scored 75.80, 68.76 and 98.61 respectively, and a 2022 follow-up in early adulthood found a foster-care group average of 73.49 against 64.65 for the group that remained institutionalized — a statistically real, roughly nine-point advantage that had held for close to twenty years. A synthesis of the full two-decade dataset, published in 2023, described the IQ effect of foster care versus institutional care as remarkably stable across the entire follow-up period.

Bar chart of average IQ by care condition at age 4.5 in the Bucharest Early Intervention Project: 109 for never-institutionalized children, 81 for children randomly assigned to foster care, and 73 for children who remained in institutional care
Bar chart of average IQ by care condition at age 4.5 in the Bucharest Early Intervention Project: 109 for never-institutionalized children, 81 for children randomly assigned to foster care, and 73 for children who remained in institutional care

The "24-month window" claim, and why it deserves a hedge

The original 2007 report, and most popular coverage of it since, highlighted a striking finding: children placed into foster care before 24 months of age recovered more fully than those placed later, suggesting a sensitive window for cognitive development. That finding is real, but the later follow-ups that tested it directly did not keep reproducing it cleanly. At age 8, the cutoff that mattered had shifted to around 26 months and applied to a narrower set of measures. At age 12, the researchers explicitly tested cutoffs at 20, 22, 24 and 26 months and found no significant effect of any of them on IQ. By early adulthood, age at placement was analyzed as a continuous variable with only a marginal association, and no discrete threshold was reported at all. The honest version: a real advantage to earlier placement shows up in this dataset, but a clean, fixed "24-month cutoff" is closer to the original headline than to what the full 20 years of follow-up data actually settled on.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

A second, independent study that points the same way

The English and Romanian Adoptees study, a separate research group following a different cohort of children adopted out of Romanian institutions into UK families, found a 15-point IQ deficit at age 11 for children institutionalized more than six months before adoption, compared with those adopted earlier — with no further worsening for longer stays between six and 42 months. Unlike the Bucharest trial, placement into adoption here was not randomized by researchers — families made their own decisions about which children to adopt and when — so this is a natural experiment rather than a controlled one. That its findings point the same direction as a true randomized trial is exactly what makes it a useful second check, not a substitute for one. A 2017 follow-up of the same cohort into young adulthood added a genuinely useful complication: the cognitive impairment measured at ages 6 to 11 had largely remitted by early adulthood, even though psychiatric and social difficulties in the same group had not. Early deprivation, in other words, did not damage every domain equally or permanently — cognition proved more recoverable over the long run than mental health did in this particular cohort.

Why toxic stress is different from an ordinary bad day

Researchers at Harvard’s Center on the Developing Child draw a specific three-way distinction: positive stress is brief and normal, like a first day of childcare; tolerable stress is serious but time-limited and buffered by a supportive relationship; toxic stress is strong, frequent or prolonged activation of the body’s stress-response systems without that buffering adult relationship in place. The proposed biological mechanism, laid out most influentially by the neuroscientist Bruce McEwen, is that repeated activation of stress hormones produces cumulative wear on brain regions, particularly the hippocampus and prefrontal cortex, that are central to memory and reasoning. That is a plausible, well-established mechanism for how chronic stress could affect cognition, not itself a specific measured effect size — the actual evidence for the size of the effect is the Bucharest and Romanian-adoptee data above.

What the ACE study does, and does not, actually show

The Adverse Childhood Experiences study, published in 1998 and often invoked in discussions like this one, is worth being precise about: it is a retrospective survey of more than 8,000 adult patients linking childhood adversity to later adult disease and risk behavior, and it never measured IQ, cognitive test performance or academic outcomes at all. Citing it as evidence that toxic stress lowers IQ specifically overstates what the study actually did. A separate, later body of work has since made that link directly: a 2024 meta-analysis pooling 32 studies and nearly 27,000 people found a real, small-to-medium association between a higher adversity score and worse cognitive control in adulthood. That is a genuine cognition finding — it is simply a different, more recent literature than the original 1998 study, which never tested cognition at all. But even one of the ACE study’s own original co-authors has since published a caution against using the ACE score for individual-level clinical prediction, and a 2021 analysis found its accuracy at predicting any single person’s outcome to be poor, even though the population-level pattern is real. Population-level evidence and individual-level prediction are different claims, and this literature is a clean example of why that distinction matters.

What this means, and does not mean

None of this says a difficult early environment determines a fixed outcome. It says something narrower and better supported: prolonged, unbuffered early adversity is one of the few environmental factors with genuine experimental, not just correlational, evidence behind its effect on cognitive development, alongside the kind of factors already covered on this site in the piece on socioeconomic status and IQ, lead exposure and IQ, and a very different prenatal exposure covered in this site’s companion piece on fetal alcohol spectrum disorder. The size of the effect is real but moderate, it responds to intervention, as the foster-care comparison shows directly, and — per the Romanian adoptee follow-up above — even the cognitive piece of it is not necessarily permanent. The single clearest, most actionable finding across all of this research is also the simplest one: the single largest protective factor identified in both the Bucharest and Romanian cohorts was the presence of one stable, responsive caregiving relationship, which is precisely the buffer the Harvard framework above names as the difference between tolerable stress and toxic stress in the first place.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged adverse childhood experiences, child iq, cognitive development, early childhood development, early intervention, foster care, institutionalization, intelligence research, IQ Science, longitudinal study, poverty and iq, randomized controlled trial, Socioeconomic Status, toxic stress

Fetal Alcohol Spectrum Disorder and IQ

Research & Evidence

Fetal Alcohol Spectrum Disorder and IQ: The Profile a Score Hides

A clinical study of 473 people diagnosed with fetal alcohol syndrome or its milder forms found that only 16 percent had an IQ in the intellectually disabled range. That statistic is not the reassurance it sounds like — it is the reason the condition is so often missed. Here is what a normal-looking score can still be hiding.

Bar chart of executive-function deficits in fetal alcohol spectrum disorder versus typically developing peers, in standard deviation units: planning 0.94, fluency 0.87, set-shifting 0.87, working memory 0.84, vigilance 0.52, inhibition 0.50 below average, from Kingdon et al. 2016

Less often than you would guess, and that gap between expectation and reality is exactly why the condition is so frequently missed. A clinic-referred study of 473 people diagnosed with fetal alcohol syndrome or a milder related condition found that only 16 percent scored in the intellectually disabled range on an IQ test. Most had an average, low-average or even above-average full-scale score — while showing a specific, measurable pattern of difficulty in a handful of other cognitive skills the score never touches.

That combination, an unremarkable composite number sitting on top of a real and specific deficit, is not a loophole in the diagnosis. It is close to the central clinical problem with fetal alcohol spectrum disorder (FASD), and it is the reason researchers have spent two decades pushing clinicians away from IQ alone and toward a specific cognitive profile instead.

What prenatal alcohol exposure does, and why there is no established safe amount

Alcohol crosses the placenta freely and can interfere with neuronal migration and brain development at multiple stages of pregnancy, with the developing brain vulnerable throughout, not only in a single early window. The CDC states plainly that "there is no known safe amount of alcohol use during pregnancy" and "no safe time during pregnancy to drink alcohol," a position the American College of Obstetricians and Gynecologists and the U.S. Surgeon General have separately reached as well. It is worth being precise about what that guidance actually claims: it is a statement that no safe threshold has been established, not a claim that every exposure at every level produces measurable harm. Heavy and frequent drinking in pregnancy is well documented to carry real, dose-related risk; the effects of very light or one-time exposure are far less firmly established either way, and treating the two claims as identical overstates what the evidence supports in either direction.

Why FASD is an umbrella, not one diagnosis

An Institute of Medicine panel in 1996 set out four categories under the FASD umbrella: fetal alcohol syndrome (FAS, the most severe, defined by a specific triad of growth deficiency, facial features and central nervous system dysfunction), partial FAS, alcohol-related neurodevelopmental disorder (ARND), and alcohol-related birth defects. The DSM-5 added a research category in 2013, neurobehavioral disorder associated with prenatal alcohol exposure, that drops the facial-feature requirement entirely and defines the condition by function instead. Worth saying plainly rather than glossing over: there is still no single international diagnostic standard. Several systems, the original IOM categories, a 2016 Canadian guideline built around impairment across specific neurodevelopmental domains, and a separately developed 4-Digit Diagnostic Code used in Washington State, currently coexist, and a 2025 systematic comparison found only fair-to-moderate agreement between them. That lack of consensus is itself a clue to what makes this condition hard to pin down.

Fewer than 1 in 10 people with an FASD diagnosis show the facial features most people associate with it. The rest, the large majority, look physically typical, which is one reason — alongside the ordinary-looking IQ scores below — that the condition is so easy to overlook entirely.

The number that surprises people

Full FAS, the most visibly affected end of the spectrum, carries a reported average full-scale IQ of around 70. The much larger nondysmorphic group, most of ARND, averages closer to 80, solidly in the low-average range and well outside the range most people picture as "intellectually disabled." The 1996 clinical study cited above found that even inside a sample of people who had already received an FAS or related diagnosis, only 16 percent scored low enough on IQ testing to meet criteria for intellectual disability. The obvious implication is the one that matters clinically: an ordinary IQ score is common in FASD, not an exception to it, and cannot be used on its own to rule the condition in or out.

Bar chart of executive-function deficits in fetal alcohol spectrum disorder versus typically developing peers, in standard deviation units: planning 0.94, fluency 0.87, set-shifting 0.87, working memory 0.84, vigilance 0.52, inhibition 0.50 below average, from Kingdon et al. 2016
Bar chart of executive-function deficits in fetal alcohol spectrum disorder versus typically developing peers, in standard deviation units: planning 0.94, fluency 0.87, set-shifting 0.87, working memory 0.84, vigilance 0.52, inhibition 0.50 below average, from Kingdon et al. 2016
Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

The profile underneath an ordinary-looking score

A 2016 meta-analysis in the Journal of Child Psychology and Psychiatry pooled the executive-function research and found a specific, consistent pattern relative to typically developing peers: large deficits in planning (0.94 standard deviations below average), verbal fluency (0.87), set-shifting (0.87) and working memory (0.84), with smaller but still real gaps in sustained attention (0.52) and inhibitory control (0.50). Other work has found working-memory deficits specifically that hold up even after global IQ is statistically controlled for — meaning the memory problem is not simply a byproduct of a lower overall score, but something that shows up on top of it. The same meta-analysis compared FASD directly against ADHD, since the two conditions are frequently confused: fluency and planning deficits were significantly worse in FASD, while vigilance and inhibition were not reliably distinguishable between the two groups, a genuinely useful and specific finding for anyone trying to tell them apart.

A complication worth stating honestly

Not every study finds executive function singled out this cleanly. A large 2021 analysis pooling six prospective U.S. cohorts, adjusted for socioeconomic status and other substance exposure, found alcohol-associated effects of roughly similar size across learning and memory, executive function, reading and math, rather than one domain standing out sharply from the rest — and found no statistically significant effect on sustained attention specifically, where the meta-analysis above did find one. The honest summary is that a specific, disproportionate executive-function signature is well supported by a substantial body of clinical research, but it is not uncontested, and the most recent large multi-cohort analysis complicates a version of the story that treats it as fully settled.

Why the diagnosis gets missed so often

A school-based study cited in a 2019 Lancet Neurology review found that fewer than 1 percent of the children who actually met FASD criteria on structured assessment had ever received a clinical diagnosis. An ordinary-looking composite score, a lack of the facial features most people expect, and a genuine lack of international diagnostic consensus among clinicians themselves all point the same direction: this is a condition that hides in plain sight far more often than it announces itself. Current U.S. prevalence estimates, from active-screening studies in four communities, run from roughly 11 to 50 per 1,000 children by a conservative count and considerably higher, 31 to 98 per 1,000, in a more thorough weighted estimate — CDC’s rounded public figure, "up to 1 in 20," is the upper end of that range, not a single settled number.

None of this is about assigning blame. Research on how this condition gets discussed has moved deliberately away from framing centered on a mother’s choices, in part because messaging that reads as accusatory has not been shown to change drinking behavior and can make people less willing to seek a diagnosis or support once a child is already showing signs. The clinical goal is identifying the actual cognitive profile early enough to help, not assigning responsibility for it after the fact.

What this means for testing and support

The practical lesson runs the same direction as it does for the conditions covered in this site’s companion piece on ADHD, autism and dyslexia: a single composite score was never built to capture a profile this specific, and for FASD in particular, a normal full-scale number is common enough that it should never be treated as evidence against the diagnosis on its own. Formal neuropsychological testing, the kind used in a full child assessment rather than a single screening number, is what actually surfaces the planning, fluency and working-memory pattern described above, and it is that pattern, not the IQ score sitting on top of it, that should drive decisions about accommodations and support.

This is one entry in a small set of ways early life can shape later cognitive outcomes through a mechanism that has nothing to do with inherited ability — see this site’s companion piece on chronic early-life stress for a very different kind of exposure that leaves a recognizably similar signature: an average score that is not the whole story. For the much broader question of how much of any IQ score is inherited in the first place, see is IQ genetic.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged child iq, cognitive development, diagnosis and assessment, executive function, FASD, fetal alcohol spectrum disorder, intelligence research, IQ Science, iq testing for children, nature versus nurture, neurodevelopmental disorder, prenatal alcohol exposure, Test Accommodations, Working Memory

High-IQ Societies Explained

Scores & Scales

Beyond Mensa: How Triple Nine, Prometheus and Other High-IQ Societies Work

Mensa is the famous one, but it is nowhere near the most exclusive. The Triple Nine Society admits one person in a thousand; the Prometheus Society, one in thirty thousand. Here is how each threshold is actually set, what test you would need, and why the further out you go, the less any single number can be trusted.

Bar chart comparing the rarity of Mensa, Triple Nine Society, Prometheus Society, and Mega Society admission thresholds, from 1 in 50 up to 1 in 1,000,000 of the population

Several societies do, and by a wide margin. Mensa’s well-known threshold is the 98th percentile — the top 2 percent of the population. The Triple Nine Society sets its bar at the 99.9th percentile, roughly one person in a thousand. The Prometheus Society goes further still, to the 99.997th percentile, about one in 30,000, and the Mega Society further still beyond that. Each step down this list is a real, order-of-magnitude jump in rarity rather than a marginal increment, and each society comes with its own quirks about which tests actually qualify a candidate for membership.

The further out this list goes, though, the shakier the underlying numbers get — not because anyone is being dishonest, but because of a limit this site has covered before: no test can reliably measure a score that rare. Every society here uses the exact same 98th-percentile logic Mensa does; what changes from one to the next is simply how far right on the curve the cut-off sits, and how much less trustworthy that cut-off becomes the further right you go.

Mensa’s threshold, as a baseline

Mensa requires the 98th percentile on an approved, supervised test of intelligence — nothing else about age, education or background counts. On the SD15 scale most modern tests use, that works out to just under 130; this site’s full breakdown of that number, and why different publishers report it as 130 or 132 covers the derivation in detail and is not repeated here. What matters for the rest of this article is the baseline: everything that follows is progressively rarer than the one high-IQ society most people have actually heard of.

Triple Nine Society: one in a thousand

The Triple Nine Society takes its name from its threshold: "999" for the 99.9th percentile, about one person in a thousand. On the SD15 scale that is 146; on SD16, 149. It accepts scores from a wide range of standardized intelligence and academic-aptitude tests — over twenty of them — which makes it considerably more accessible than either of the two societies further down this list. Not easy: a 1-in-1,000 score is still rare enough that the overwhelming majority of people who sit a standard test will never see one on their own report, no matter how well they happen to do. But it is reachable through an ordinary supervised assessment, the same kind of test a school, clinic or employer might administer, rather than a specialist instrument built specifically to chase a rarer tail of the distribution.

Prometheus Society: one in thirty thousand

The Prometheus Society sets its bar at the 99.997th percentile — 160 on the SD15 scale, 164 on SD16, and roughly one person in 30,000. Here the practical requirements change in a way that is worth understanding rather than just noting: as of today, the society accepts essentially only the Miller Analogies Test (MAT) for qualification. That is not an arbitrary rule. Most standardized IQ tests simply are not built with enough discriminating power that far out in the distribution to certify a score that rare in the first place — which is the same ceiling problem covered in detail in why no IQ test can reliably score you above about 160.

Bar chart comparing the rarity of Mensa, Triple Nine Society, Prometheus Society, and Mega Society admission thresholds, from 1 in 50 up to 1 in 1,000,000 of the population
Bar chart comparing the rarity of Mensa, Triple Nine Society, Prometheus Society, and Mega Society admission thresholds, from 1 in 50 up to 1 in 1,000,000 of the population
Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

Mega Society: one in a million, and a cautionary history

The Mega Society sits past Prometheus again, at the 99.9999th percentile — about one person in a million. Founded in 1982 by Ronald K. Hoeflin, it is a useful case study in what goes wrong at this level of rarity, and it is not merely a matter of statistics. Hoeflin’s own qualifying instrument, the Mega Test, ran unsupervised and untimed in Omni magazine in 1985 — anyone could take it at home, on paper, at their own pace. That design made sense for reaching a geographically scattered population of extremely rare scorers, and it was also its downfall: once enough people had discussed the questions and shared answers, later scores stopped meaning what earlier ones had. The society stopped accepting Mega Test scores from after 1994, and a successor instrument, the Titan Test, was eventually retired for the same reason in 2020.

That history matters beyond one society’s administrative footnote: it is a second, independent reason — alongside the statistical extrapolation problem below — that scores this far into the tail are hard to trust at face value. A test rare and specialised enough to target one-in-a-million scorers is also, almost by construction, too small and too specialised in its niche to be administered under the same tightly controlled, regularly re-normed conditions as a mainstream instrument like the WAIS.

Why the numbers get shakier as the club gets more exclusive

This is the important part, and it follows directly from that ceiling-effect article: a standardisation sample for a major test typically runs into the low thousands of people. That is plenty to characterise the middle of the distribution precisely, and it is nowhere near enough to empirically verify what a 1-in-30,000 or 1-in-a-million score should actually look like. Past a certain point, a test manual is extrapolating the shape of a curve outward from data it does not really have, not reading a measurement off a ruler. That is exactly why the Mega Society, at the 99.9999th percentile — one person in a million — is not given a specific point-score equivalent here. Past a certain rarity, a single number claims a precision the underlying test data cannot support.

What these societies actually do

Beyond the admission threshold, the day-to-day reality of Triple Nine, Prometheus and similar societies looks a lot like Mensa’s: member newsletters, online discussion groups, occasional in-person meetups, and special-interest groups built around members’ other hobbies. The exclusivity is entirely in the door, not in what happens after you are through it — joining one of these does not unlock anything functionally different from what Mensa already offers, just a smaller and more specifically selected room.

That is worth sitting with for a moment, because it cuts against the mystique these societies sometimes attract online. There is no secret curriculum, no exclusive research programme and no functional advantage to membership beyond the social one of meeting other people who cleared the same unusually high bar. For most people genuinely curious about their own score, an ordinary supervised test and a clear explanation of where the result sits on a standard scale will answer the actual question far more usefully than pursuing admission to any of these.

Should you try to qualify

If you are curious, the realistic first step is the same one either way: sit a properly normed, supervised test and see where you land, rather than assuming an online quiz score tells you anything about eligibility for any of these. A casual, ungraded online test is not accepted by any of these societies, and for the same reason it should not be treated as a real estimate of where you would land on one that is — the ceiling and norming problems described above apply just as much to an unverified quiz as they do to a specialist admission test, only with far less rigor behind the number it reports. And if you do not clear one of the higher bars, that is at least as likely to reflect the ceiling problem described above as it is to reflect anything meaningful about you — the tests that can distinguish the 98th percentile reliably were mostly never built to distinguish the 99.997th. Worth remembering, too: none of this measures the kind of skill covered in this site’s piece on memory versus IQ — these thresholds are about reasoning under test conditions, not how much you can memorize. For your own number, on a test that reports the scale, percentile and confidence range together rather than a bare figure, start here.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged admission requirements, cut-off scores, extreme scores, gifted, high iq societies, iq percentile, iq scale, IQ Science, mega society, mensa, prometheus society, rare scores, standard deviation, triple nine society

Memory and IQ

Understanding IQ

Does a Great Memory Mean You’re Smart? Memory and IQ, Explained

The world’s best memory athletes can memorize 71 shuffled words in twenty minutes. Their IQ scores are not measurably different from a matched control group who could only manage 40. If a spectacular memory is not the same thing as a high IQ, what actually is the difference?

Two-panel bar chart comparing memory athletes and matched controls: 71 versus 40 words recalled after 20 minutes, but nearly identical fluid-reasoning scores of 128.1 versus 128.4

Not especially, no — and the cleanest evidence for that comes from a study that took the question as literally as possible. Researchers recruited 23 of the world’s top memory athletes, people who can memorize a shuffled deck of cards in under a minute, and compared them against 23 ordinary people matched for age, sex and IQ. On memory tasks, the athletes were not close — they recalled 71 of 72 words after a 20-minute delay against 40 for the controls. On a standard measure of fluid reasoning, the two groups scored 128.1 and 128.4. Statistically indistinguishable.

That gap between a massive skill difference and a nonexistent IQ difference is the whole story here, and it says something specific about what IQ tests measure that most people get wrong on their own — including, often, people who assume a sharp memory for names, dates or trivia is itself a sign of a high IQ.

The study that tested this directly

Dresler and colleagues, publishing in the journal Neuron in 2017, gathered 23 of the world’s top-50-ranked competitive memorizers and compared them, brain scans included, against 23 controls deliberately matched for age, sex, handedness and IQ — recruited partly from Mensa and academic-foundation mailing lists specifically so the comparison group would already be drawn from a high-functioning population, not an average one. That design choice matters: it means this study cannot tell you whether memory athletes outscore the general population on IQ (they likely do, since both groups here were well above average). It can tell you something more useful — whether the specific skill of extreme memorization tracks with reasoning ability even among people who already reason well. It does not.

Competitive memory itself is a real, standardized sport with its own disciplines: memorizing the order of a shuffled deck of playing cards as fast as possible, recalling long strings of random binary digits, matching dozens of unfamiliar faces to names seen once, and reciting memorized decimal digits of pi. None of these events resemble an IQ test’s matrix-reasoning or analogy items even slightly — they reward a specific, trainable encoding skill applied to arbitrary material, not the ability to spot a novel pattern you have never been drilled on.

A huge gap on one measure, none on the other

After studying a list of 72 words for a fixed period, memory athletes recalled an average of 71 of them 20 minutes later. Matched controls recalled 40 — a real skill, just not a remotely comparable one. On the fluid-reasoning measure, though, the two groups landed at 128.1 and 128.4 respectively, a difference small enough to be noise. The athletes were dramatically better at one very specific thing and not measurably better at the kind of on-the-spot problem-solving an IQ test is built to capture.

Two-panel bar chart comparing memory athletes and matched controls: 71 versus 40 words recalled after 20 minutes, but nearly identical fluid-reasoning scores of 128.1 versus 128.4
Two-panel bar chart comparing memory athletes and matched controls: 71 versus 40 words recalled after 20 minutes, but nearly identical fluid-reasoning scores of 128.1 versus 128.4

What made the athletes better, if not IQ

Brain imaging pointed to connectivity, not raw processing power: memory athletes showed stronger functional connections between brain networks involved in memory and in visual processing, consistent with the technique nearly every competitive memorizer uses, the method of loci — mentally placing items to be remembered along a familiar route and "walking" that route to retrieve them. The most striking part of the study is what happened when naive controls were taught this technique and given six weeks to practise: many approached the athletes’ recall performance, closing a gap that had looked biological. That result belongs to a pattern seen again and again across skill research, and this site has covered it before — see how much of expert performance generally comes down to technique and deliberate practice rather than a fixed, innate ceiling.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

Does training your memory make you smarter

This study is also a useful, if indirect, data point against a much broader claim: that training one narrow cognitive skill hard enough produces general gains in intelligence. The memory athletes in this study have trained one specific skill for years, often at a serious competitive level, and it produced exactly the outcome you would predict if training were narrow rather than general — a huge improvement on the trained skill and no detectable improvement on fluid reasoning. Commercial "brain training" programmes make a version of the opposite promise, and the research on whether narrow drills transfer to general intelligence is genuinely mixed at best; this study is a fairly clean illustration of why skepticism is the reasonable default position to start from.

Why IQ tests barely touch this kind of memory

IQ tests do include a memory component, but it is a narrower one than most people picture: working memory, typically tested with tasks like repeating a string of digits forwards and then backwards, or holding a small amount of information in mind while manipulating it. That is genuinely a different cognitive skill from what a memory athlete trains, which is long-term, associative memory for large amounts of arbitrary material — a skill IQ tests were never designed to assess, because it responds so strongly to technique and practice rather than to the kind of novel, on-the-spot reasoning the tests are built around. For more on why even working memory itself is not the same thing as reasoning, see this site’s companion piece on the news desk, and for the IQ-test component that measures how fast you process information rather than how much you can hold, see processing speed and IQ.

The everyday version of this confusion

The same mix-up shows up outside competitive memorizing. "Photographic memory" as popularly imagined — a perfect, effortless mental snapshot of anything seen once — has never been reliably documented in a controlled study of an adult. What looks like it in daily life is usually domain-specific expertise: chess masters can reconstruct a mid-game board position from a glance not because their memory is generically superior, but because years of practice let them recognize and "chunk" familiar patterns instead of memorizing 32 pieces one at a time. This was demonstrated decades earlier, well before the memory-athlete study, in a classic experiment by Chase and Simon: put those same chess masters in front of a board with pieces placed randomly, in configurations that never occur in real play, and their recall advantage over a novice mostly disappears — because there are no real game patterns left to recognize. The skill was pattern recognition trained on a specific domain, not raw memory capacity, and it evaporates the moment the domain-specific structure is removed. The same logic applies to a memory athlete: ask one to recall a list of 72 unrelated, un-memorable historical dates instead of 72 concrete words their method of loci was built to handle, and the advantage would very plausibly shrink for exactly the same underlying reason.

None of this is a knock on the athletes. Encoding 72 words in a way that survives 20 minutes is a genuinely difficult, learnable skill, and they are simply better at it than almost anyone alive.

What this means if you are prepping for an IQ test

Mnemonic tricks are worth learning for their own sake, but they will not meaningfully move a real IQ test score, because the tests are deliberately built around novel problems you have not seen before and cannot pre-memorize a solution to. A digit-span task, which does appear on most IQ tests as part of the working-memory component, is one of the few places a memorized technique could theoretically help a little — but it is a small slice of the overall score, and the matrix-reasoning and verbal-analogy sections that carry most of the weight are immune to memorization by design. For what actually helps in the days before sitting one, see this site’s guide to preparing for an IQ test. And if raw memory is not what separates an ordinary score from an exceptional one, the next question is what does — which is exactly where the numbers get genuinely interesting: at the far, rarely-discussed end of the scale, covered in a look at what sits above Mensa.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged brain training, cognitive ability, fluid intelligence, general intelligence, intelligence testing, iq myths, IQ Science, memory, memory athletes, memory champions, memory techniques, method of loci, mnemonics, Working Memory

Preschool and IQ

Research & Evidence

Does Preschool Raise a Child’s IQ? What the Perry Project Found

The most famous experiment in early-childhood education raised participants’ IQ scores by as much as 13 points — then watched that advantage shrink to almost nothing within a few years. Their income, education and arrest records stayed different for the rest of their lives. Why would outcomes outlast the score that supposedly predicts them?

Bar chart comparing Perry Preschool program and no-program groups at age 40: 65 versus 45 percent high school graduation, 64 versus 45 percent avoiding five or more arrests, and 49 versus 15 percent earning over 20,000 dollars, despite the early IQ advantage fading by age eight

Yes, in the short term, and the size of that boost has been measured about as rigorously as anything in this field: a randomized trial that raised participants’ IQ scores by roughly 13 points. The complication is what happened next. Within a few years the IQ advantage had shrunk to almost nothing. What did not shrink was everything the IQ score was supposedly there to predict.

That combination — the measurable thing fading while the outcomes it was supposed to forecast kept diverging for decades — is the most-cited puzzle in early-childhood research, and it changes what a preschool IQ study can actually tell you. It also means the honest answer to "does preschool raise IQ" depends entirely on which year you ask the question, which is not the kind of nuance that fits in a single headline number.

The experiment that settled the design question

The HighScope Perry Preschool Study began in 1962 in Ypsilanti, Michigan, with 123 children living in poverty, ages 3 and 4, individually assigned to either a high-quality preschool programme or no preschool at all — not assigned by classroom or neighbourhood, which is what makes the results a genuine causal estimate rather than a comparison of whoever happened to enroll. Researchers then did something almost nobody else in this field has managed: they kept following the same 123 people all the way to age 40.

What the programme actually consisted of

It was not a light touch. Children attended a daily 2.5-hour classroom session every weekday morning, taught by certified public-school teachers holding at least a bachelor’s degree, at an unusually low child-to-teacher ratio of about 6 to 1. The classroom followed what became the HighScope curriculum’s "plan-do-review" routine: children planned an activity, carried it out, then reviewed what happened, rather than sitting through passive instruction. Teachers also made a 1.5-hour home visit to each family every week, working directly with mothers on extending the same activities at home. That combination of classroom intensity, low ratios, a specific pedagogy, and weekly direct parent involvement is a considerably heavier intervention than what most children who attend some form of preschool today actually receive, which matters for how far the results can be expected to generalise.

The IQ boost, and how fast it faded

At age 5, right after the programme ended, the treatment group’s IQ advantage measured about 0.75 standard deviations — on some analyses as much as 13 points, a large effect by any standard in psychology. By age 8, that gap had shrunk to roughly 0.08 standard deviations, about a single point, and was no longer statistically distinguishable from zero. The children who attended preschool were, by nearly any IQ test given in elementary school, no longer measurably different from the children who had not.

Bar chart comparing Perry Preschool program and no-program groups at age 40: 65 versus 45 percent high school graduation, 64 versus 45 percent avoiding five or more arrests, and 49 versus 15 percent earning over 20,000 dollars, despite the early IQ advantage fading by age eight
Bar chart comparing Perry Preschool program and no-program groups at age 40: 65 versus 45 percent high school graduation, 64 versus 45 percent avoiding five or more arrests, and 49 versus 15 percent earning over 20,000 dollars, despite the early IQ advantage fading by age eight

The outcomes that did not fade

Here is where the study becomes genuinely strange if you expect IQ to be doing the predictive work. Followed to age 40, the preschool group graduated high school at 65 percent versus 45 percent for the no-preschool group. They were arrested five or more times over their lives at 36 percent versus 55 percent. They were earning $20,000 or more annually at 49 percent versus 15 percent. Researchers estimate the programme’s impact on lifetime earnings at upward of $200,000 per participant. Every one of those gaps persisted for decades after the IQ gap that supposedly explained early advantage had already closed in elementary school.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

Why would outcomes outlast the IQ boost

The leading explanation is not that the IQ tests were wrong, but that IQ was never the actual mechanism of benefit — it was just the easiest thing to measure early. The economist James Heckman and colleagues, reanalyzing this same dataset, argue the programme’s lasting effect ran mainly through non-cognitive skills: self-control, sustained attention, motivation and the ability to work with other children — skills that do not show up on an IQ test at all but that plausibly shape whether a teenager stays in school or a young adult keeps a job. A more recent reanalysis (Garcia, Heckman and colleagues, 2025) goes further, arguing some of the apparent cognitive fadeout was itself a measurement artefact of how later tests were scored, and that real cognitive gains persisted further into adulthood than the original age-8 result suggested. That reanalysis is contested, not settled — but even by the original, more conservative reading, the outcome gap needs an explanation that is not simply "the IQ gap stayed open."

One further piece of evidence favours the soft-skills reading over a purely cognitive one: through elementary and middle school, years after the IQ gap itself had closed, the preschool group was still less likely to be placed in special education and less likely to be held back a grade than the no-preschool group. If the programme’s only lasting effect had been cognitive, those school-progress differences should have closed on the same schedule as the IQ scores. They did not, which is exactly the pattern you would expect if the programme had changed something more durable than a test score — how a child navigated a classroom, not just how they scored on one test.

Head Start’s larger, messier evidence

Perry Preschool is a single, small, unusually intensive programme from the 1960s, run at a scale and cost per child that no national system has matched, and its dramatic numbers do not automatically generalize to preschool at national scale. Head Start, the actual federal early-childhood programme serving over a million American children a year, has been studied far more broadly, and its own large-scale evaluation (the national Head Start Impact Study) found a similar shape — an early cognitive boost that fades within a few years of school entry — but with smaller, more debated effects on later life outcomes than Perry’s striking numbers. The honest summary is that intensive, well-resourced early intervention for children in poverty has a real and reasonably well-established effect on adult outcomes; exactly how large that effect is at the scale of an ordinary national programme is a genuinely open, actively studied question, not a settled multiple of the Perry numbers.

How this differs from does school raise IQ

This is a narrower question than whether ordinary schooling raises IQ across the school years, which looks at the general K-12 population. This article is scoped specifically to targeted early-childhood interventions for children growing up in poverty, where the comparison is not "more school versus less school" but "an intensive structured programme versus none, before age 5." The two questions share a family resemblance and not much more.

It also rhymes, more than coincidentally, with the one randomized trial on breastfeeding and IQ: another case where a real, causally established early cognitive advantage shrinks well before adulthood. The through-line in both is the same caution — an IQ score measured in early childhood is a snapshot of that moment, not a fixed prediction of the adult the child will become.

What this means, and does not mean

None of this makes IQ scores meaningless, and it does not mean preschool is pointless — if anything, Perry Preschool is some of the strongest causal evidence available that early intervention for children growing up in poverty changes the trajectory of a life. What it means is narrower and more useful: a single IQ number, especially one measured in early childhood, is an incomplete predictor of how a life turns out, because plenty of what shapes an adult outcome was never going to show up on that test in the first place.

It is also a useful corrective for reading any "programme X raised IQ by Y points" headline, in education research or anywhere else: ask what happened to that gap five years later before deciding how much weight it deserves. Sometimes, as with Perry Preschool, the score itself was never really the point.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged cognitive development, early childhood education, early intervention, head start, HighScope, iq fadeout effect, IQ Science, longitudinal study, Perry Preschool, poverty and iq, preschool, randomized trial, school readiness, Socioeconomic Status