The sturdiest results are narrow and decades old. The 2025 studies that drew the headlines rest on self-reports and small preprints — and none of them measured intelligence at all.
AI assistants have been everyday tools for a little under four years. The research on what they do to human thinking is only now arriving, and it says less than the headlines built on top of it. The sturdier findings are narrow: short-term, task-specific effects on memory and on effortful reasoning. The loudest ones tend to come from one-time surveys, which cannot establish cause and effect at all.
Researchers file most of this work under one unglamorous term: cognitive offloading. It means pushing mental work onto something outside your head — a notepad, a calendar, a colleague, a search box, a language model. The behavior is ancient. What is new is the scale, the speed, and the fact that the external helper now hands back finished reasoning rather than raw ingredients.
Psychologists were studying offloading long before generative AI existed. A much-cited 2011 experiment published in Science reported that when people expect information to be saved somewhere they can reach later, they remember the content itself less well and remember where to find it better. That is a trade, not a straight loss. It also comes with a caution. Parts of this literature have a mixed replication record — meaning that when other labs ran the same experiments again, they sometimes found smaller or less consistent effects. Treat any single study here as one data point, not a verdict.
The cleanest everyday example is navigation. Across a range of studies, people who follow turn-by-turn directions learn the layout of a route less well than people who work it out from a map. Both groups arrive. The difference is what is left behind: nothing in the guided task required building a mental map, so no mental map got built. You get the outcome without the by-product. That is the pattern offloading research keeps turning up.
The AI-specific wave of 2025 is thinner than its coverage suggests. One widely cited study, by researchers at Microsoft Research and Carnegie Mellon University, surveyed knowledge workers about how they used generative AI in real work. Those who expressed more confidence in the tool also reported putting less critical effort into the task; those with more confidence in their own expertise reported putting in more. Two limits matter. The measure is self-report, so it captures what people believe about their own thinking rather than how they actually performed. And the design is cross-sectional — a single snapshot of many different people, rather than the same people followed over time — which cannot separate whether the tool reduced effort from whether people already inclined to spend less effort reach for the tool more often.
A second study travelled further. An MIT Media Lab team had a few dozen participants write essays in one of three conditions — with a language model, with a search engine, or unaided — while wearing EEG caps, scalp sensors that pick up electrical activity from the brain. The assisted group showed weaker measures of coordinated activity across brain regions during writing, and had noticeably more trouble quoting their own essays back shortly afterwards. The authors listed the caveats themselves: few participants, one narrow writing task, and release as a preprint — a paper posted publicly before peer review, the process in which independent specialists check the methods and can force corrections. One more caveat belongs alongside those. EEG does not read thinking directly. It records electrical signals that researchers then interpret, and the interpretations are contested.
The reliable finding is not that people became less intelligent. It is that work you never did yourself leaves fewer traces behind.
No published study shows that using AI assistants lowers general intelligence. Demonstrating that would require the same people measured on the same properly normed test — one first given to a large, representative sample so that an individual result can be placed on a standard scale — before and after years of heavy use, alongside a comparison group who used the tools less. No such dataset exists, because mainstream AI assistants are less than four years old and studies of that shape run for many years. “AI is making us dumber” is, for now, an extrapolation rather than a finding.
The other evidence often dragged into the argument is the so-called reverse Flynn effect. Through most of the 20th century, average raw performance on IQ tests rose from generation to generation — the Flynn effect. Because publishers re-norm their tests periodically, resetting the average to 100 each time, that rise is visible only in the raw scores. Since roughly the 1990s, several national datasets have shown the climb flattening, and dipping on some subtests. Two things are worth noting. In several countries the turning point comes before smartphones existed at all. And the candidate explanations — schooling, nutrition, test-taking motivation, immigration and shifts in who gets sampled — are still actively argued over. It is not a clean AI story.
The most useful framing to come out of this research is a split most coverage collapses. Assisted performance — what you can produce with the assistant open — is clearly rising. Unassisted capability — what you can do when it is closed — is what these studies gesture at, and it is measured far less often. Both are real, and they can move in opposite directions at once without any contradiction. The gap only becomes visible when the tool is absent: a closed-book exam, an outage, a live conversation, a judgment call where you have to decide whether the confident answer on your screen is actually right.
If you want a read on where your unaided reasoning stands, the measurement has to be careful enough to mean anything. A properly normed test places your result on a named scale — most often a mean of 100 with a standard deviation of 15, which in plain words means roughly two in three people score between 85 and 115 on that scale. A percentile describes the share of the reference group scoring below you, and it is an estimate, not an exact rank. A single sitting produces an estimate rather than a fixed property of you, so an honest result comes with a confidence interval: a range several points wide within which your true score most likely sits. A bare number with no named scale and no range is not a measurement. That standard matters more, not less, in a period when people are asking whether their own thinking has changed.
Find your IQ score now! →The honest summary today is narrow and slightly boring. Offloading mental work to machines reliably changes what you retain from that particular work — a pattern shown for decades with notebooks, calculators and maps, and now being shown with language models. Whether a few years of that adds up to a durable change in general reasoning ability is an open question, and the evidence needed to answer it does not exist yet. Anyone telling you otherwise, in either direction, is ahead of the data.
No published study shows that. Demonstrating it would mean measuring the same people on the same properly normed test before and after years of heavy use, alongside a comparison group who used the tools less. Those studies do not exist yet, because mainstream AI assistants are only a few years old. What the research does show is narrower: information and reasoning you hand to an external tool tend to be remembered less well afterwards.
It is any use of something outside your head to do mental work for you — a shopping list, a calendar reminder, a calculator, satellite navigation, a search engine or an AI assistant. It is normal, and frequently the right choice. The research question is not whether to offload, but which mental work is worth keeping in your own head.
Not on its own. Average raw test performance rose for most of the 20th century and has flattened or dipped on some subtests in several countries since roughly the 1990s. In several of those datasets the turning point comes before smartphones existed, and the proposed causes — schooling, test motivation, sampling and population changes — are still debated. It is not a settled technology story.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
Two people can perform identically on two different online tests and walk away with scores twenty points apart. The scale, the comparison group and the margin of error are what separate an assessment from a quiz.
Average scores climbed for decades, then flattened and in places slipped. None of it shows on a score report, because every test is reset so the average is 100 again — which is why a 1990 score is not a 2026 score.
IQ scales are defined by a human reference sample, and a language model is not in it. Placing a machine on that scale is not a hard measurement problem — it is a category error.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →