IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

Does Using AI Make You Dumber?

Mind & Everyday Life

Does Using AI Make You Dumber? What the Studies Measured

The claim that AI is making us stupider has real research behind it and is still routinely overstated. No study has measured a drop in intelligence. What has been measured is lower mental engagement during a delegated task and weaker memory for work you did not do yourself, which is a narrower finding and a more useful one.

Chart of what the cognitive debt essay-writing study measured, showing lower brain engagement and weaker recall in the AI-assisted group, alongside the outcomes the study did not measure including IQ

Does using AI make you dumber? On the evidence available in 2026, no study has shown that AI use lowers intelligence, and none has measured an IQ score before and after. What researchers have measured is narrower and still worth taking seriously: people who delegate a thinking task to a language model engage less while doing it, remember less about it afterwards, and evaluate the output less critically the more they trust the tool.

Those are real findings about attention, memory and judgment. They are not findings about general intelligence, and the distance between the two is where most of the alarming coverage lives.

What the cognitive debt study actually measured

The study driving most of the headlines is work from the MIT Media Lab titled "Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task". Fifty-four adults wrote essays under one of three conditions: with a language model, with a search engine, or with no tools at all. The researchers recorded brain electrical activity during the task and analysed the resulting text.

The model-assisted group showed measurably lower cognitive engagement while writing. They also showed weaker recall of their own essays afterwards — in several cases struggling to quote work they had submitted minutes earlier. The authors called the accumulated effect cognitive debt: a shortfall you take on by skipping the effortful part, which comes due later.

  • The finding is about engagement during the task and memory for the output.
  • It is not a finding about reasoning ability, problem solving, or any score on a cognitive test.
  • Nobody in the study was given an intelligence test at any point.

What the study did not show

Being clear about the limits is not a way of dismissing the work. It is how the work should be read, and the authors are considerably more careful than the coverage.

  • Fifty-four people is a small sample. Effects this size in samples this size routinely shrink when a study is repeated at scale.
  • One task, one session. Essay writing under observation over a short window is not the same as habitual use over years, which is what the headline claim implies.
  • Brain engagement is not intelligence. Lower measured activity during a task you delegated is close to what you would predict. It does not follow that capacity changed.
  • Cognitive debt is a metaphor, coined by the authors and not a validated psychological construct with an established measure behind it.

The correct summary is that a well-designed small study found lower engagement and weaker recall under AI assistance, and that this is a reason to pay attention rather than a demonstration that anyone got less intelligent. The site applies the same standard to claims that things raise intelligence, which usually turn out to be weaker than advertised in exactly the same way.

Chart of what the cognitive debt essay-writing study measured, showing lower brain engagement and weaker recall in the AI-assisted group, alongside the outcomes the study did not measure including IQ
Chart of what the cognitive debt essay-writing study measured, showing lower brain engagement and weaker recall in the AI-assisted group, alongside the outcomes the study did not measure including IQ

The finding that should worry you more

A separate line of research points at something more specific than general dulling. A survey by researchers at Microsoft and Carnegie Mellon University found that the people who most trusted the accuracy of AI assistants thought least critically about what those assistants produced. Confidence in the tool, rather than time spent with it, predicted the drop in scrutiny.

Work tracking professional consultants found the same shape from the other direction: measurable short-term performance gains, reported in the range of 14 to 40 percent, alongside erosion of the independent judgment the assistance was supposed to support. The mechanism is not mysterious. Automation handles the routine cases and hands you the exceptions, which removes exactly the ordinary practice that keeps judgment sharp — so when the tool is wrong, the person checking it is out of practice.

Judgment behaves like a skill rather than a trait. It decays without use, and it does not show up on an intelligence test either way, which is part of why smart people can be reliably bad at particular kinds of decision. We look at that gap in why high scorers still make poor decisions.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

What would a study need to show to settle this?

It is worth being concrete about the gap between what exists and what would actually answer the question, because the gap is large and nothing currently published closes it.

  • A normed cognitive measure, before and after. Not brain activity during a task, but a standard instrument administered at the start and again at the end.
  • Random assignment and a real control group. People who choose to use AI heavily differ from people who do not, in ways that would produce this pattern on their own.
  • Months or years, not one session. The claim is about habitual use, so the study has to run long enough for a habit to exist.
  • A correction for practice effects. Sitting the same test twice raises the second score whatever happened in between, and that alone can manufacture or mask an effect.

No published study does all four. That is not a criticism of the researchers — a study like that is expensive and slow — but it does mean anyone claiming AI has measurably lowered human intelligence is going beyond the evidence. The same standard applies to the reverse claim: nobody has shown it is harmless either.

Does this apply to children and students?

The research discussed here was conducted on adults, and extending it to developing minds is exactly the kind of leap the evidence does not support. Children are not small adults for these purposes: the skills at issue are still being acquired rather than maintained, which could plausibly make delegation more costly or less, and neither has been demonstrated.

What is reasonably well established is narrower and older: learning that involves retrieval and effortful practice sticks better than learning that does not. A tool that removes the effort removes the thing that made it stick. That is an argument about how the tool is used in teaching, not about whether it lowers intelligence, and it long predates language models.

Is this the same as cognitive offloading?

Related but not identical, and the distinction is worth keeping. Cognitive offloading is the long-studied habit of storing information outside your head — a phone number in a contacts list, a route in a map app — and the research on what it does to memory predates language models by decades. We cover that literature separately in our report on AI and memory.

The cognitive debt work is about something narrower: not what happens to your memory when you store a fact elsewhere, but what happens to your engagement and recall when you delegate the thinking itself. Offloading a phone number costs you the number. Offloading the reasoning may cost you the practice.

How to use AI without losing the practice

Nothing in this research supports avoiding these tools, and the productivity findings are as real as the engagement ones. What the evidence does support is being deliberate about which part of the work you hand over.

  • Attempt first, then delegate. Producing your own answer before asking for one preserves the effortful step the studies found missing, and gives you something to compare against.
  • Use it to critique rather than to produce. Asking a model to find the weakness in your reasoning keeps you doing the reasoning.
  • Distrust fluent output on purpose. The Microsoft and Carnegie Mellon result says confidence in the tool is the risk factor, so the correction is to check most carefully when the answer reads best.
  • Keep some work unassisted. Not for virtue: for the same reason anyone practices anything they intend to stay good at.

If the underlying worry is about your own thinking rather than the technology, the useful move is to measure rather than speculate. A properly normed reasoning test gives you a baseline you can compare against later, reported as a position in a reference population rather than a bare number. And if the deeper question is whether machines are overtaking us, the domain-by-domain comparison is a better guide than any single study.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged AI and intelligence, artificial intelligence, brain health, chatgpt, cognitive debt, cognitive offloading, critical thinking, does ai make you dumber, IQ Science, learning, mental effort, sample size, study quality, Working Memory

What Is ChatGPT’s IQ Score?

Scores & Scales

ChatGPT IQ Score: Why the Numbers Disagree

Search for ChatGPT’s IQ score and you will find 155, 136, 116 and under 100, all reported seriously, sometimes for the same model. They are not contradictions to be resolved. They are different tests, scored different ways, and the spread between them is the most informative thing about the whole exercise.

Chart of reported AI IQ figures from four different tests, ranging from below 100 on a culture-fair test to 155 on the verbal subtests of the Wechsler scale, showing how far the same systems spread across tests

There is no single ChatGPT IQ score. Published figures for the same family of systems range from slightly below 100 on one culture-fair test to 155 on the verbal half of a clinical instrument, with 116 and 136 reported in between. Every one of those numbers was produced honestly. They disagree because they come from different tests, administered under different conditions, and because none of them is an IQ in the sense the word has when a psychologist uses it.

That spread is worth understanding, because the same forces distort the score a person gets from an online test.

What IQ scores has ChatGPT actually been given?

Four results account for most of what circulates, and they are not measuring comparable things.

  • 155 on the Wechsler verbal subtests. A clinical psychologist administered parts of the Wechsler Adult Intelligence Scale and reported a verbal IQ of 155, above 99.9 percent of the American standardisation sample. Only the verbal subtests could be given: the performance subtests need eyes, ears and hands.
  • Under 100 on a culture-fair test. On the Worldwide IQ Test, a non-verbal assessment designed to minimise language and cultural loading, GPT-4o scored slightly below the population average of 100.
  • 116 and 136 on two other tests. OpenAI’s o3 model was reported at 116 on one assessment and at 136 on the public Mensa Norway test, which would place it above roughly 98 percent of people if it were a person.
  • Up to 151 on tracked weekly testing. The TrackingAI project has been administering the Mensa Norway test to frontier models every week; by September 2026 the top of that chart had reached 151.

A system cannot be simultaneously in the top 0.1 percent and below average. What varies is the test.

Why does the same system score 155 and under 100?

The Wechsler verbal subtests reward stored knowledge: vocabulary, general information, verbal similarities. That is the closest thing to a language model’s home ground, and the result reflects it. The culture-fair test does the opposite. It strips out language and prior knowledge deliberately and asks for pattern completion on abstract figures, which is precisely the ability these systems have found hardest.

In a person, this rarely happens, because human abilities correlate: someone with a top-percentile vocabulary usually does well on matrices too. In a machine there is no such tie between the two, so the choice of test decides the answer. The same principle explains why verbal and non-verbal index scores can diverge sharply in a human profile too, and why a single full-scale number can hide it.

Chart of reported AI IQ figures from four different tests, ranging from below 100 on a culture-fair test to 155 on the verbal subtests of the Wechsler scale, showing how far the same systems spread across tests
Chart of reported AI IQ figures from four different tests, ranging from below 100 on a culture-fair test to 155 on the verbal subtests of the Wechsler scale, showing how far the same systems spread across tests

How do you even give an IQ test to a chatbot?

Less straightforwardly than the headline numbers suggest, and the administration details do a lot of work. The Mensa Norway test behind most published AI figures is a set of 35 visual-pattern puzzles. A model that cannot see is given those puzzles described in words; a model that can see is given the original images. TrackingAI runs both variants weekly and reports the average of the last seven administrations rather than a single sitting.

That produces something genuinely instructive. The same underlying system often appears twice on the chart, once reading a description and once looking at the picture, and the two scores are not the same. One model scored 133 when the puzzles were described to it and 136 when it could see them. Another scored 113 reading and 108 looking. The direction is not even consistent.

  • The format of administration moved the score by several points without anything about the system changing.
  • Averaging seven sittings hides how much any single sitting varies, which for a person is exactly what a confidence interval is for.
  • A test built to be visual becomes a different test when it is read aloud, in the same way a timed test becomes a different test when the clock is removed.

Human testing treats this as fundamental rather than incidental. Standardised administration — same instructions, same time limit, same materials — is part of the instrument, because two scores only mean the same thing if they were produced the same way. Almost none of the AI figures in circulation were.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

What happens when the test has never been published?

This is the most revealing comparison available, and it comes from the same project that produces the headline numbers. TrackingAI runs two tests on every model it follows. One is Mensa Norway, a public online test that has been on the internet for years and is therefore in the training data. The other is an offline test written by a Mensa member that, in the project’s own words, has never been on the public internet and is in no AI training data.

Reading the underlying chart data on 5 September 2026, 27 models carried a score on both tests. Across those 27, scores on the public test averaged about 13.5 points higher than scores on the unpublished one — close to a full standard deviation on the usual IQ scale. The public test topped out at 151; the unpublished test at 136.

  • One widely used model scored 137 on the public test and 113 on the unpublished one, a 24-point gap.
  • Another scored 145 publicly and 127 privately.
  • A third scored 108 publicly and 88 privately.
  • A small number moved the other way, which is a useful reminder that this is a pattern rather than a law.

An honest caveat belongs here, because it is the kind this site insists on. Two different tests are not directly comparable: they have different norming samples, different difficulty and different item formats, and the unpublished test’s construction has not been released. The gap is consistent with the public test having leaked into the training data, but it does not prove it on its own. What can be said plainly is that models score substantially higher on the test they have almost certainly seen.

The human parallel is exact, and it is the reason serious tests are kept out of circulation. Someone who has worked through a test before scores higher on it the second time without having become any smarter. That is the practice effect, and it is a measured, predictable thing rather than a suspicion.

Is an AI IQ score a real IQ?

No. An IQ is not a mark out of anything. It is a position within a reference population, expressed on a scale with a stated mean — almost always 100 — and a stated standard deviation, usually 15. Producing one requires a norming study in which the test is given to a representative sample of that population under standard conditions.

For a machine, there is no population to be a member of and no norming sample, so the final step simply cannot be taken. What a model produces on an IQ test is a count of correct answers. Calling that count an IQ borrows the authority of a scale it was never placed on. We make the full argument in our piece on why a model scoring 120 is not scoring an IQ.

There is a second problem specific to machines. Standard administration assumes a fixed time, no external help and no second attempt. A model may be run at different effort settings, with or without tools, once or many times, and the reported figure is usually the best of those runs. A human score reported that way would not be accepted either.

What this means for your own score

The lesson transfers directly. A number without a stated scale and a stated reference group is not a result, whoever produced it. When you take a test online, the questions that matter are which population your score is being compared against, what the scale’s standard deviation is, and how much measurement error sits around the figure — a point we cover in the explainer on why a score is a range, not a number.

It also means treating a score you got on a test you had already seen with the same scepticism you would apply to a model’s 151. If you want a number that means something, take a properly normed reasoning test once, cold, and read the result as a range rather than a point. For what the resulting figure does and does not tell you, how to read a test report is the place to start, and the wider human-versus-machine comparison puts the AI numbers in context.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged AI and intelligence, AI benchmarks, artificial intelligence, chatgpt, chatgpt iq score, Culture Fair Test, intelligence test, iq scale, IQ Score, mensa, norming sample, standard deviation, test norms, wais

Is AI Smarter Than Humans?

Understanding IQ

Is AI Smarter Than Humans? What the 2026 Benchmarks Show

Ask whether AI is smarter than humans and the honest answer depends entirely on the task. Frontier systems now beat expert humans on graduate-level science questions and lose to ordinary people on puzzles built to be unfamiliar. Here is what the 2026 benchmarks actually measure, and why one score never settles it.

Chart comparing machine and human performance across four task types, showing machines ahead on graduate science questions and coding, level on novel puzzle solving, and behind on tasks never seen before

Is AI smarter than humans? On narrow, well-specified tasks the answer in 2026 is often yes, and increasingly by a wide margin. On problems a system has never seen before, described by nobody, with no worked example to copy, people still hold the advantage. The reason both statements are true at once is that "smarter" is not one axis, and the tests that make machines look unbeatable and the tests that stop them cold are measuring genuinely different things.

This article takes the comparison domain by domain, using figures published by the labs and by independent evaluators in 2026, and then explains the structural difference that makes a single answer impossible.

What AI already does better than most people

The clearest machine wins are on tasks with a correct answer, a large body of prior examples, and no requirement to act in the world. On GPQA Diamond, a set of graduate-level biology, chemistry and physics questions written to be hard for people with access to a search engine, OpenAI reported GPT-6 Astra at 96.0 percent in September 2026. Gemini 3.1 Pro sits at 94.3 percent. Both are above the performance of domain experts answering outside their own speciality.

The same pattern holds across coding, long-document retrieval and structured professional work. These are not trick results. They are real, they are reproducible, and they describe abilities that took people years of training to acquire.

  • Recall and synthesis at volume. No person holds the contents of a technical literature in working memory. A model effectively does.
  • Speed. Work that takes an expert a day is returned in minutes, which changes what is worth attempting.
  • Consistency. A model does not get tired on the four-hundredth item, which is exactly where human scorers drift.
  • Breadth of surface knowledge. Competence across far more fields than any individual can maintain.

Where humans still hold the edge

The sharpest counterexample is ARC-AGI-3, a benchmark from the ARC Prize Foundation built specifically to test learning rather than recall. A system is dropped into a small turn-based environment with no instructions and has to work out the goal, the controls and the rules by acting inside it. Crucially, every environment is calibrated on people first: humans solve 100 percent of them, because a task is only admitted once people have shown it can be done.

That calibration is what makes the comparison meaningful, and it is the detail most coverage drops. We set out the full argument in our report on what ARC-AGI-3 measures. The short version: when the novelty is real and the instructions are absent, the gap between an ordinary adult and a frontier system has been enormous.

That gap is now closing fast, and the way it closed is instructive. In September 2026 the ARC Prize Foundation reported GPT-6 Astra at 62.7 percent on its neutral, provider-independent test harness — and 99.9 percent on a harness supplied by the model’s own developer, which preserves the system’s reasoning state between moves. Same model, same benchmark, same week. The scaffolding around the model accounted for most of the difference.

Is AI smarter than humans at learning new things?

This is the question the benchmark was built to ask, and the 2026 answer is genuinely mixed. On action efficiency — how many moves a solver needs to work an unfamiliar environment out — the ARC Prize Foundation found that Astra used fewer actions than the median tested human on 96 percent of the levels it completed, and 51.7 percent fewer actions per level on average. By that measure the machine matched and passed human parity.

Cost tells a different story. The human baseline came from around 500 members of the public, paid roughly 12.78 dollars per game attempted. The model runs that produced those scores cost between 17,332 and 26,098 dollars. The system that learns as efficiently as a person in moves does so at several thousand times the price in resources.

Both numbers are real, and neither alone answers the headline question. That is the pattern to expect from here.

Chart comparing machine and human performance across four task types, showing machines ahead on graduate science questions and coding, level on novel puzzle solving, and behind on tasks never seen before
Chart comparing machine and human performance across four task types, showing machines ahead on graduate science questions and coding, level on novel puzzle solving, and behind on tasks never seen before
Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

Why smarter is the wrong question

In people, mental abilities correlate. Someone who scores well on vocabulary tends to score well on spatial reasoning and on arithmetic, and the pattern is consistent enough that a single summary number carries real information. That pattern is the reason IQ works at all — we explain the underlying statistics in our explainer on the g factor.

Machines do not show that structure. A system can answer graduate physics correctly and then fail a coloured-grid puzzle that a ten-year-old solves in thirty seconds. Its abilities do not hang together, so no single number summarises them, and any ranking against a person depends entirely on which task you picked. This is not a temporary measurement problem. It is a real difference in how the two kinds of system are built.

  • Human ability is correlated; one score generalises across many tasks.
  • Machine ability is jagged; world-class in one domain, below average in the next, with no reliable pattern.
  • So a comparison needs a task, and any claim without one is not a measurement.

Does a benchmark score mean the same as an IQ score?

No, and the distinction matters more than it sounds. A benchmark result is a raw percentage: items solved out of items attempted. An IQ is not a percentage of anything. It is a position within a reference population, expressed on a scale with a defined mean, almost always 100, and a defined standard deviation, usually 15.

Turning a raw count into that position requires a norming study: the same test, administered under standard conditions, to a representative sample of the population the score will be read against. No such sample exists for machines, and it is not clear what one would even be. So a model can score 96 percent on a science exam without that number converting into any IQ at all. We take that argument apart properly in the article on what ChatGPT’s IQ score really is.

The same logic governs human scores, which is why a raw count on a reasoning test is meaningless until it is placed against a reference sample. That placement is exactly what the IQ percentile calculator does, and it is the step a benchmark percentage has no equivalent for.

When will AI be smarter than humans overall?

Predictions from serious people vary by decades, which is itself the most useful fact about them. Geoffrey Hinton has said he expects machines to surpass human intelligence within about twenty years. Others working on the same systems put it sooner or reject the framing entirely. There is no measurement that would settle the disagreement, because there is no agreed test — the problem we work through in the explainer on what AGI actually means.

What can be said with confidence is narrower. The hardest evaluations still defeat the best systems: on Humanity’s Last Exam, a set of expert-written questions across dozens of fields, the strongest reported result in September 2026 was 65.0 percent, meaning better than a third of the questions remained unanswered. Benchmarks that were supposed to hold for years keep falling, and new ones keep being built because the old ones stop separating anything.

What this means for measuring your own intelligence

None of this changes what a cognitive test does for a person. The reason matrix puzzles sit near the centre of most non-verbal reasoning tests is that they lean as little as possible on what you happen to know and as much as possible on working out a rule from the evidence in front of you. That the same format is what machines found hardest is a point in favour of the format, not against it.

If the comparison has made you curious about your own reasoning rather than a model’s, our IQ test is built around exactly this kind of rule-finding, and the result is reported the way a score has to be reported to mean anything: as a position in a reference population, with a range around it. For what those numbers do and do not predict, the evidence on outcomes is a better guide than any headline about machines.

The durable conclusion is unglamorous. Machines are now better than most people at a growing list of specific things, worse at a shrinking list, and not comparable at all on the single scale the question implies. Anyone offering you one number for it is selling something.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged AI and intelligence, AI benchmarks, artificial intelligence, cognitive ability, fluid reasoning, G Factor, general intelligence, human intelligence, intelligence research, intelligence test, IQ Science, is ai smarter than humans, machine learning, reasoning tests

Does Exercise Raise Your IQ?

Mind & Everyday Life

Does Exercise Raise Your IQ? What Aerobic Fitness Changes

Aerobic exercise has some of the best-replicated cognitive evidence of any lifestyle factor, but not for a full IQ score. What reliably improves is executive function: working memory, inhibitory control, cognitive flexibility. One related skill, planning, barely moves at all.

Bar chart of aerobic exercise effects on four executive function skills, showing real improvements in inhibitory control, working memory and cognitive flexibility, and no significant change in planning

For a specific, well-defined slice of cognition — called executive function, covering things like ignoring distraction, holding information in mind, and switching between tasks — yes: aerobic exercise has a real, repeatedly replicated benefit. For a full IQ score, the case is thinner, and unlike chess, video games or reading, exercise is not really a story about practicing a mental skill at all. It works, as far as anyone can tell, through the body.

That is a genuinely different mechanism from the other three activities in this batch, worth sitting with before comparing the results. Chess and reading are both about practicing something mentally specific and asking how far the benefit travels. Exercise is a physiological intervention that happens to have cognitive side effects, and the research asks a slightly different question: how much exercise, for whom, delivered how, and sustained for how long, rather than how many hours of deliberate practice a particular skill needs.

What moves, and by how much

Meta-analyses of aerobic-exercise programs in healthy middle-aged and older adults find real improvements on three separate executive-function measures: cognitive flexibility, working memory, and inhibitory control — the ability to hold back an automatic but wrong response. All three effects are small to moderate in size and consistent enough across studies to be taken seriously.

A fourth measure, planning ability, did not show a statistically reliable improvement in the same body of research. That is worth including precisely because it is the negative result: exercise is not simply a uniform boost to every kind of executive skill, and a chart that only showed the three positive findings would be more flattering than accurate.

  • Cognitive flexibility. A real, moderate improvement — the largest of the four.
  • Working memory. A real, moderate improvement, close behind flexibility.
  • Inhibitory control. A real but smaller improvement.
  • Planning. No statistically reliable improvement in this population.
Bar chart of aerobic exercise effects on four executive function skills, showing real improvements in inhibitory control, working memory and cognitive flexibility, and no significant change in planning
Bar chart of aerobic exercise effects on four executive function skills, showing real improvements in inhibitory control, working memory and cognitive flexibility, and no significant change in planning

The same pattern in children

The adult numbers above are not an isolated finding. Separate reviews of aerobic-exercise programs in children and adolescents with ADHD found moderate improvements across the same three skills — inhibitory control, working memory and cognitive flexibility — and a review focused on overweight and obese children found a similar moderate improvement in overall executive function, driven mainly by inhibitory control and working memory rather than cognitive flexibility.

Lining those studies up next to each other matters more than any single number in them. Three reviews, three different populations — healthy older adults, children with ADHD, children carrying excess weight — using different exercise programs and different research teams, and all three land on roughly the same short list of skills: inhibitory control and working memory move fairly reliably, cognitive flexibility often does too, and the least consistent finding across all of them is planning. A single study finding a benefit could easily be a fluke. The same shape of result recurring across unrelated populations is a much harder pattern to explain away.

Why aerobic exercise might do this

The proposed mechanisms behind this are physiological rather than purely psychological in nature: sustained aerobic activity increases blood flow to the brain and appears to raise levels of growth factors involved in forming new connections between neurons, particularly in regions tied to memory and executive control. Human evidence for the exact chain from a training program to a measured cognitive gain is still being worked out, and much of what is confirmed comes from a mix of animal research and indirect markers in humans rather than one complete human causal pathway. The consistent behavioral result — better executive function after aerobic training — is on firmer ground than any single explanation for why it happens.

There is also a difference between a single session and a sustained program, and the two do not always point the same way. A single bout of moderate aerobic exercise can produce a short-lived boost to attention and inhibitory control lasting roughly an hour or two afterward, which researchers call an acute effect. A program of regular exercise sustained over weeks or months produces the more durable executive-function gains described above, through what looks like a slower, structural change rather than a temporary state. The two are related but not interchangeable, and a study measuring one says little about the other.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

Dose matters, and so does your starting point

The size of the benefit depends on how much exercise, how intense, and for how long, and studies that vary these do not all agree on the exact best combination. What is more consistent across this research is who benefits most: people who start out less fit, or with lower baseline cognitive performance, tend to show larger gains than people who were already fit and high-performing. The same general pattern — real benefits in children with attention difficulties, in sedentary older adults, and in children carrying excess weight — shows up across strikingly different groups, which is itself a form of evidence: a result this consistent across such different starting populations is harder to explain away as a fluke of any one study design.

That “lower baseline benefits more” pattern also shows up when comparing a single exercise session’s after-effects directly: people who perform worse on a cognitive task before exercising tend to show a bigger improvement afterward than people who already scored well. It is a pattern with an intuitive ceiling built in — there is more room for someone starting lower to move — but it also means the people most often used to headline this research, young healthy volunteers already near their own ceiling, may be exactly the group least likely to show a dramatic effect.

Executive function is not the same as a full IQ score

It matters that almost none of this research measures a full-scale IQ score directly. Executive function is one contributor to test performance, particularly on timed and working-memory-heavy sections, but a standard IQ battery also leans on things exercise research rarely touches, like accumulated vocabulary and abstract pattern reasoning. Processing speed and working memory are the parts of a typical test most plausibly connected to what exercise studies actually measure; treating a gain on an executive-function task as equivalent to a higher IQ score overstates what the research supports.

This is also why exercise and diet sit oddly next to chess, video games and reading in one respect: none of the exercise research reviewed here reports a full-scale IQ score before and after a program, the way a handful of the chess and brain-training studies at least attempt to. What exists is evidence about specific, named cognitive skills that overlap with part of what an IQ test measures, not a demonstrated change in the composite score itself. That is a genuine gap in the evidence, not a minor technicality, and a fair summary of this research has to say so rather than round “executive function improved” up to “IQ went up”.

So should you exercise for a sharper mind

For executive function specifically, aerobic exercise has some of the best-replicated evidence in this entire batch of questions, consistent across ages from childhood through later life; see how cognitive performance shifts across the lifespan and what early scores predict about later-life health for the surrounding picture. For a general IQ score, treat exercise the way this site treats diet: a real contributor among several, not a single lever. Nutrition and this article cover two of the more physiological factors; the full picture of what does and does not move the number ties them together.

The version of this claim worth actually believing is the modest one: regular aerobic activity is one of the better-supported ways to keep specific, useful mental skills sharp, particularly for anyone starting from a lower baseline, and that is true regardless of what it does or does not do to a single test score.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged aerobic exercise, BDNF, brain health, cognitive development, cognitive flexibility, executive function, exercise and iq, how to increase iq, Improving IQ Scores, intelligence test, IQ Science, nature vs nurture, physical activity, Working Memory

Does Reading Raise Your IQ?

Mind & Everyday Life

Does Reading Raise Your IQ? The Matthew Effect, Explained

Reading volume and vocabulary growth are genuinely linked, and the relationship compounds over years rather than staying fixed. But it is mostly the verbal, crystallized side of ability that moves, not the abstract reasoning an IQ test also measures, and cause and effect run both ways.

Line chart of vocabulary size across school grades for three reading-volume groups, starting close together in early grades and fanning out into a wide gap by the later grades

Reading volume and vocabulary growth are genuinely, strongly linked, and the relationship compounds over years rather than staying flat — psychologists call it a Matthew effect, after the biblical line about the rich getting richer. What it mostly moves, though, is verbal ability: vocabulary, background knowledge, comprehension. It is weaker evidence that reading raises the broader, more abstract reasoning an IQ test also tries to capture, and untangling cause from effect turns out to be harder than the popular version of this claim admits.

That distinction — which kind of ability actually moves — is the same question this batch keeps returning to with video games and exercise: a real, well-documented effect on something specific, next to a much weaker claim about intelligence in general.

The Matthew effect, named for a very old line

The psychologist Keith Stanovich described the pattern in the 1980s: children who decode text easily read more, and reading more builds vocabulary and background knowledge, which makes the next book easier to read, which leads to reading still more. Children who struggle with decoding do the opposite — they read less, encounter fewer new words, and fall further behind readers who started out only slightly ahead of them. The gap is not there from day one. It opens gradually, driven by a feedback loop rather than a single cause.

The name comes from a line in the Gospel of Matthew about the rich getting richer, and it is now used across psychology and economics for any process where a small early advantage compounds into a large later one, rather than staying the size it started at.

The same shape of feedback loop turns up outside reading too — in wealth, in athletic training, in scientific reputation — anywhere a small early edge changes how much opportunity or practice follows. Reading is simply one of the better-studied examples, because schools measure both sides of the loop, reading skill and vocabulary, on the same children year after year.

What the numbers show

This is not just a plausible story; it shows up in longitudinal data. One study tracking children from kindergarten through the later grades found that word-reading skill in fourth grade predicted the rate of vocabulary growth afterward, not just the vocabulary a child already had — and this held up even after statistically accounting for how large a child’s vocabulary already was back in kindergarten. Above-average readers were not just ahead; they kept pulling further ahead.

The same body of research found first-grade reading ability predicting outcomes measured in eleventh grade, a full decade later, and that the prediction survived even after removing the part explained by earlier general cognitive-ability scores. In plain terms: how well a child was reading in first grade told researchers something real about where they would land by the end of school, beyond what an early IQ-type score alone would have predicted.

Line chart of vocabulary size across school grades for three reading-volume groups, starting close together in early grades and fanning out into a wide gap by the later grades
Line chart of vocabulary size across school grades for three reading-volume groups, starting close together in early grades and fanning out into a wide gap by the later grades

How researchers try to separate cause from effect

Simply asking children how much they read is a weak measure, because struggling and confident readers describe their own habits very differently. A workaround used across much of this literature is a print-exposure checklist: a long list of real book and author titles mixed in with invented ones that sound plausible, where the score is how many real titles a person recognizes. It is a rough proxy for how much a person has actually read over the years, and because it does not ask anyone to self-report or take a vocabulary test directly, it gives researchers a way to measure reading volume that is not simply the same thing as the vocabulary score it is being used to predict.

Twin studies add a second angle. Comparing identical and fraternal twins raised in the same household lets researchers estimate how much of the overlap between reading habits and vocabulary is really about shared genes and shared upbringing, rather than reading causing vocabulary directly. That work generally finds a real, independent contribution from reading itself — it is not purely a proxy for something else the twins already had in common — but it is a smaller contribution than the raw, unadjusted correlation between reading and vocabulary would suggest on its own.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

Which kind of smarter this is actually about

Psychologists split intelligence into a fluid part — reasoning through a genuinely new problem with no learned content to draw on — and a crystallized part, built from accumulated knowledge and vocabulary. The Matthew-effect research is almost entirely about the second kind. Verbal IQ scores draw heavily on vocabulary and general knowledge, which is exactly what wider reading builds most directly; the more abstract, pattern-based reasoning on the non-verbal side of most tests has a much thinner connection to how many books someone has read. This split between accumulated, knowledge-based ability and raw, in-the-moment reasoning is one of the oldest and best-replicated distinctions in the field; other frameworks for describing distinct mental abilities cover related ground from a different angle.

Cause and effect also run in both directions, which the phrase “reading raises your IQ” tends to flatten into one. Children with larger early vocabularies find reading easier and therefore do more of it; the reading then builds the vocabulary further. Some of what looks like reading’s effect is really an early head start showing up again later, and the honest summary is a loop that reinforces an early difference rather than a one-way lever anyone can pull from a standing start.

Reading, EQ and other kinds of ability

It is worth being precise about what widening vocabulary actually buys a reader, socially as well as academically. A larger vocabulary and more background knowledge make it easier to follow, produce and be persuaded by complex arguments, which is a real advantage in school and at work even before it shows up as a higher test score. It is a different kind of advantage from the interpersonal skills covered in IQ versus EQ, which draw on a mostly separate set of abilities that reading volume on its own does not obviously move.

Does this still apply once you are an adult

Most of the strongest evidence here comes from childhood and the school years, when vocabulary and background knowledge are being built fastest and the gap has the most time to compound. The case for adult reading habits moving a fully developed vocabulary by a similar mechanism is much thinner — not because it has been disproven, but because it has simply been studied far less. What adult reading almost certainly still does is maintain and extend specific knowledge and vocabulary in whatever a person reads about, which matters for real-world communication and comprehension even without any change to a test score.

There is also a practical difference in what “reading” means at each age. A school-age Matthew effect is mostly about whether a child reads at all, and how much, since the comparison is against children who barely read outside class. An adult who already reads fluently is instead varying the topic, difficulty and volume of material that is, relatively speaking, a much smaller manipulation — more like choosing a harder workout than starting to exercise from nothing. That difference alone would predict a smaller effect in adulthood even if the underlying mechanism never changed.

So does reading raise your IQ

For the crystallized, vocabulary-and-knowledge side of ability, the evidence that reading volume matters is some of the strongest in this whole batch of questions — stronger than the case for chess or video games moving anything at all. For the fluid, reasoning side that IQ tests also measure, the case is much weaker, and at least part of the childhood effect is an early difference compounding rather than reading creating an advantage from nothing. The broader question of what actually moves an IQ score covers where reading fits next to environment and the other factors this site has looked at.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged cognitive development, crystallized intelligence, how to increase iq, Improving IQ Scores, intelligence test, IQ Score, Matthew effect, nature vs nurture, print exposure, reading and iq, reading habits, verbal intelligence, vocabulary growth