Galton’s Victorian measuring instruments. Binet’s first real test. Wechsler’s bell curve. A century and a half of figuring out how to put a number on a mind — and what that number can honestly claim to mean. This is the full story, from the first attempts to measure intelligence through reaction time, to the adaptive tests scored in your browser today.
For most of human history, mental ability was judged by reputation, apprenticeship or birthright, not by a number. One of the earliest attempts to rank people by intellectual merit rather than social class was China’s imperial examination system, the keju, used from roughly the 7th century onward to select civil servants. It tested classical literary and administrative knowledge rather than intelligence in the modern sense, but it established a principle intelligence testing would later borrow: standardized questions, administered the same way to everyone, used to make consequential decisions about people’s lives.
The birth of psychology as an experimental science is usually dated to 1879, when Wilhelm Wundt opened the first dedicated psychology laboratory in Leipzig, Germany. One of his American students, James McKeen Cattell, coined the term “mental test” in 1890 and — influenced heavily by Francis Galton — tried to measure intelligence through simple sensory and motor tasks: reaction time, grip strength, sensitivity to pain, the ability to judge short time intervals. It was a reasonable hypothesis: keener senses, sharper mind. In 1901, Cattell’s own student Clark Wissler tested it directly, correlating these measurements against university students’ actual grades. The correlations came back close to zero. Sensory-motor speed, it turned out, had almost nothing to do with academic performance — a negative result that quietly cleared the way for a very different approach a few years later in France.
China’s keju system selects civil servants by standardized written examination rather than birthright — an early precedent for testing, not intelligence testing itself.
Wilhelm Wundt opens the first dedicated experimental psychology laboratory in Leipzig, founding psychology as a measurable science.
Clark Wissler shows Cattell’s sensory “mental tests” don’t correlate with school grades — clearing the way for Binet’s approach.
Sir Francis Galton (1822–1911) — a British polymath and half-cousin of Charles Darwin — was arguably the first person to treat intelligence as something that could be measured, quantified and compared across a population. In his 1869 book Hereditary Genius, Galton argued that intellectual ability was primarily inherited, based on his observation that eminence tended to cluster within certain families.
In 1884 Galton opened an Anthropometric Laboratory in London, where visitors paid a small fee to have their reaction time, grip strength, head size and sensory acuity measured — assembling data on more than 9,000 people. His most durable contribution, though, wasn’t any single test; it was statistical. Working with the tools that became correlation and regression analysis, later formalized by his student Karl Pearson, Galton gave psychology a language for describing how strongly two measurements move together — the mathematical backbone every intelligence test still relies on today. In 1904, British psychologist Charles Spearman built directly on this foundation, proposing that performance on many different mental tasks correlated because they all drew on one underlying general ability, which he labeled simply “g.”
See how modern norming builds on this statistical foundation →Binet built a tool to identify children who needed help — not a permanent verdict on how intelligent they were.
A French government request for a practical classroom tool became the template every intelligence test since has followed.
In 1904, the French Ministry of Public Instruction asked psychologist Alfred Binet to solve a practical problem: identify which children in Paris’s public schools were likely to struggle with a standard curriculum, early enough to give them extra help. Binet and his collaborator Theodore Simon rejected Galton and Cattell’s sensory approach entirely. Instead, the Binet-Simon Scale, published in 1905, tested judgment, comprehension, reasoning and vocabulary directly: following simple commands, naming objects in pictures, repeating a string of digits, explaining how two objects differ, defining abstract words. Many of these item formats are still recognizable in intelligence tests today, including the reasoning sections of our own Verbal IQ Test and Numerical IQ Test.
The 1905 scale’s most influential idea was mental age: a child who answered questions as well as the average 8-year-old was assigned a mental age of 8, regardless of their actual birth date. Revisions in 1908 and 1911 organized test items by age level and extended the scale into adolescence. Because children develop at uneven rates, comparing mental age to chronological age gave educators a simple, practical way to flag a meaningful gap — precisely the tool the French government had asked for. Our own Kids IQ Test still leans on the same core principle: age-appropriate norming rather than one fixed adult standard.
Binet himself was notably cautious about how far this idea should be pushed. He warned against what he called the “brutal pessimism” of treating a test score as a fixed ceiling on a child’s ability or worth, insisting the scale was meant to identify who needed help, not to rank children permanently by native intelligence. It’s a caution the next chapter of this history did not always heed.
Binet’s test found only modest interest in France, but it crossed the Atlantic quickly. Henry H. Goddard, director of the Vineland Training School in New Jersey, translated the Binet-Simon Scale into English in 1908 and began using it to assess residents of his institution. Goddard was also a committed hereditarian: his 1912 book The Kallikak Family used one extended family’s history to argue that “feeble-mindedness” was simply inherited, and he oversaw intelligence testing of immigrants at Ellis Island that was used to support restrictive, discriminatory immigration decisions. Most historians of psychology now treat this as one of the field’s clearest cautionary chapters — a genuinely useful diagnostic tool, applied confidently to conclusions its own data couldn’t support, with real harm to real people.
Lewis Terman, a psychologist at Stanford University, took the same raw material in a very different direction. His 1916 revision — renamed the Stanford-Binet Intelligence Scale for his university — extended and re-normed Binet’s age-based items for an American population and adopted a formula German psychologist William Stern had proposed in 1912: intelligence quotient, or IQ, calculated as mental age divided by chronological age, multiplied by 100. A 10-year-old performing at a 12-year-old’s level scored 120; one performing at an 8-year-old’s level scored 80. It was simple, it produced one clean number, and it made “IQ” — not “mental age” — the term the public would use for the next century.
Goddard’s English translation brings the Binet-Simon Scale to the United States for the first time.
William Stern proposes mental age divided by chronological age, multiplied by 100.
Terman’s revision popularizes the term “IQ” and re-norms the scale for American test-takers.
Six decades separate Galton’s measuring instruments from Wechsler’s deviation IQ — and nearly as many separate Wechsler from the adaptive, computer-scored tests in use today.
Figure 1. Key milestones in the history of intelligence testing, 1869 to today. See how today’s tests build on this lineage in our IQ Test Guide →
Every intelligence test up to this point had been administered one-on-one: an examiner sitting across from a single child or adult, working through the scale by hand. That changed in 1917, when the United States entered World War I and needed to sort more than a million new recruits quickly. A committee of psychologists chaired by Robert Yerkes — with Terman among its members — designed two tests that could be given to a roomful of soldiers at once: the Army Alpha, a written test for English-literate recruits, and the Army Beta, a non-verbal, picture- and symbol-based version for recruits who were illiterate or did not speak English.
By the war’s end, the Army had tested roughly 1.75 million men, using scores to help guide officer selection, job placement and, in some cases, discharge. It was the first large-scale demonstration that intelligence testing could work as a fast, standardized, group-administered instrument rather than a slow individual clinical exercise — a shift the field never really reversed. Our own IQ Tests hub, timed and delivered the same way to every test-taker, descends directly from that innovation.
The Army’s testing program had a second, quieter legacy. In 1926, the College Board introduced the Scholastic Aptitude Test, designed largely by Carl Brigham, a psychologist who had worked on the Army program and adapted its item formats for college admissions. The SAT was explicitly framed as a way to identify talented students who hadn’t attended elite preparatory schools — though, like Goddard’s work with Binet’s test, some of Brigham’s own early writing about group differences reflected the era’s hereditarian biases, which he later publicly repudiated in 1930 once further research undercut his original claims.
Terman’s ratio formula worked reasonably well for children, but it had an obvious flaw for adults: mental age effectively plateaus in the late teens, which meant the same adult’s IQ score would mathematically drift downward every year simply by getting older, with no real change in ability. A 30-year-old reasoning as well as an average 20-year-old would score a mere 66 under the ratio formula — clearly not what “IQ” was supposed to measure.
David Wechsler, chief psychologist at Bellevue Hospital in New York City, solved this in 1939 with the Wechsler-Bellevue Intelligence Scale by abandoning the ratio formula altogether. His deviation IQ compared a person’s raw score only to other people of the same age, then mapped that comparison onto a normal distribution — a bell curve — with a mean of 100 and a standard deviation of 15. This is still, in essence, how every major intelligence test, including ours, calculates a score today. Our IQ Bell Curve page breaks down exactly how a raw score becomes a percentile under this system, and IQ and Age shows how age-based norm groups work in practice.
Wechsler’s scales introduced a second lasting change: instead of one blended score, they reported separate Verbal and Performance IQ scores from distinct groups of subtests — an early version of the domain-by-domain breakdown (verbal, numerical, spatial, reasoning speed) that shows up in modern tests, including the format-specific assessments in our IQ Tests hub. Wechsler later built age-specific versions of this scale: the WISC (Wechsler Intelligence Scale for Children, 1949) and the WAIS (Wechsler Adult Intelligence Scale, 1955), both still in clinical use today in revised editions.
Figure 2. An illustrative comparison, not derived from real test data: the same hypothetical person’s score under the old ratio formula versus Wechsler’s deviation IQ. Because mental age plateaus in the late teens, the ratio formula scores this person lower every year simply for aging — the exact flaw deviation IQ was designed to remove. See how real age-based norming works on our IQ and Age →
If test norms were never updated, scores would drift upward for a strange reason: people, on average, keep getting better at IQ tests. Named for researcher James Flynn, who documented the pattern across dozens of countries in the 1980s, the Flynn Effect describes a rise in raw intelligence test scores of roughly three points per decade through most of the 20th century — driven by better nutrition, expanded schooling, smaller families and a more visually complex, test-like modern environment.
This is precisely why Wechsler’s deviation-IQ system needed a mechanism its 1939 version didn’t fully anticipate: periodic renorming. Every major test is re-standardized against a fresh sample every decade or so specifically to keep the mean anchored at 100 — without that step, an unchanged test would slowly report a rising average score that reflected the test going stale, not humanity getting smarter. Since the 1990s, several developed countries have reported the gains flattening or slightly reversing, an unresolved pattern researchers sometimes call the “negative Flynn Effect.” Our Can You Improve Your IQ? page goes deeper into what does and doesn’t move an individual score, and Average IQ by Country shows how much scores vary across different norm populations today.
Roughly the average rise in raw IQ scores documented across much of the 20th century.
James Flynn documents the rise across more than a dozen nations, giving the effect its name.
Several developed countries report the gains flattening, or slightly reversing — still an active research question.
The core idea hasn’t changed since Binet. What changed is the theory behind what’s being measured, and the technology used to measure it.
Every intelligence test since Binet shares the same basic recipe: give people a standardized set of reasoning problems, compare their performance to a well-defined norm group, and report the result on a common scale. Most contemporary psychologists now organize cognitive abilities using the Cattell–Horn–Carroll (CHC) model, which integrates Raymond Cattell’s split between fluid intelligence — raw reasoning with unfamiliar problems — and crystallized intelligence — accumulated knowledge — with John Carroll’s broader hierarchy of specific cognitive abilities.
The mid-20th century also produced a direct response to intelligence testing’s early cultural-bias problem. In 1936, British psychologist John C. Raven introduced the Progressive Matrices, a test built entirely from abstract visual patterns, with no reliance on language, vocabulary or culturally specific knowledge. That same instinct — measure reasoning while minimizing the influence of language, education or cultural background — is exactly what our own Culture-Fair IQ Test is built around.
The most recent shift has been delivery, not theory. Computer-adaptive testing, where the difficulty of each question adjusts in real time based on previous answers, lets a test estimate ability precisely in a fraction of the items a fixed-form paper test would need, and lets results be scored and returned in minutes instead of weeks. That’s the format behind every assessment in our own IQ Tests hub, including the Classical, Verbal, Numerical, Spatial and Timed formats — each normed the way Wechsler’s, not Binet’s, generation of psychologists established.
Early tests often assumed shared language and cultural reference points test-takers didn’t have. Nonverbal, culture-reduced formats like Raven’s Matrices were built as a direct response.
Goddard’s Ellis Island testing and Brigham’s early Army-data writing used intelligence scores to support discriminatory policy — conclusions the data never supported, and ones Brigham himself later retracted.
A score describes performance on a specific set of reasoning tasks, on a specific day — not a person’s total worth or fixed potential. The exact caution Binet himself raised in 1905.
A 1995 task force convened by the American Psychological Association, chaired by Ulric Neisser, remains the field’s most-cited attempt to separate what the research does and doesn’t support — concluding that IQ tests reliably measure something real and practically useful, while cautioning strongly against using scores to make claims about groups rather than individuals. That balance — a genuinely useful instrument, used carefully — is the thread connecting Binet’s 1905 caution to how a responsible test should be built today. See our IQ Test Accuracy page for what modern reliability research actually shows.
For a more narrative walk through this same material — including how each era’s test actually looked to the people taking it — see our companion article, History of IQ Testing: Origins to Significance, or browse ongoing research breakdowns on the IQ Metrics Blog.
To see exactly how a modern test translates everything on this page into a single score, read How IQ Tests Work, then compare formats on the IQ Tests hub before you take one yourself.
The French psychologist Alfred Binet, working with Theodore Simon, published the first practical intelligence test in 1905 — commissioned by the French government to identify schoolchildren who needed extra academic support. See how the modern descendants of that idea work in our IQ Test Guide.
IQ stands for Intelligence Quotient. German psychologist William Stern proposed the underlying ratio formula — mental age ÷ chronological age × 100 — in 1912, and Lewis Terman adopted and popularized the term “IQ” three years later in his 1916 Stanford-Binet Intelligence Scale.
The ratio formula (mental age ÷ chronological age × 100) breaks down in adulthood, because mental age effectively stops rising after the late teens — a fully-formed adult scoring like an average 20-year-old would be assigned a falling IQ every year they aged, regardless of any real change in ability. David Wechsler replaced it in 1939 with the deviation IQ, which compares a person’s score to others in their own age group. See exactly how that scale works on our IQ Bell Curve page.
During World War I, a committee of psychologists led by Robert Yerkes developed the Army Alpha, for English-literate recruits, and Army Beta, a non-verbal version for non-English speakers and illiterate recruits, to help the U.S. Army classify around 1.75 million soldiers. It was the first time intelligence testing was administered to large groups at once rather than one person at a time, and it directly shaped the group-testing format later used by the IQ Tests hub.
The Flynn Effect, named for researcher James Flynn, is the documented rise in average raw IQ test scores across generations through most of the 20th century — roughly three points per decade in many countries. It’s a major reason tests are periodically renormed against a fresh sample. Some countries have shown a plateau or slight reversal since the 1990s. Read more in Can You Improve Your IQ?
Not in their original form. Their core ideas survive — age-based norming, a mix of verbal, numerical and spatial reasoning items, group administration — but modern tests like the ones in our IQ Tests hub use updated, extensively re-normed items, computer-adaptive delivery, and statistical methods that didn’t exist in the early 20th century.
No, and this is a genuine part of the field’s history. Early 20th-century testing was sometimes misapplied to support discriminatory immigration and eugenics policy — most infamously by Henry Goddard. Modern test design responds directly to those failures, including nonverbal, reduced-language formats like our Culture-Fair IQ Test, built specifically to lower the influence of cultural and linguistic background on a reasoning score.
The reasoning being measured is more similar than most people expect — pattern recognition, verbal reasoning, working memory, numerical and spatial logic. What changed is delivery and scoring: today’s tests are self-administered online rather than one-on-one with an examiner, scored instantly against large digital norm samples instead of a printed age table, and built from item-response statistics that didn’t exist in 1905. You can see the current version of that lineage in our own IQ Test.
Every test in our lineup descends from the same research this page just walked through — scored the way Wechsler intended, delivered the way the 21st century expects.
Take the Classical IQ Test →