Analyses that compare siblings inside the same family do find a firstborn advantage on measured intelligence — roughly one to two points on a mean-100, standard-deviation-15 scale. It is real, its size is disputed, and it is no larger than the error on a single test.
Of all the folk theories about intelligence, birth order has the strongest claim to being partly true. Ask whether the eldest child is smarter and the honest answer is not no. Studies that compare siblings raised in the same house do find a firstborn advantage on measured intelligence. The advantage is on the order of one to two points on a scale where the average is 100 and the standard deviation is 15 — small enough to be invisible inside any actual family, no larger than the error attached to a single test score, and nothing like the difference the popular version implies.
Start with the scale, because a number without one means nothing. Modern intelligence tests are scored so that the population average is 100 and the standard deviation — a measure of how far apart people's results typically fall — is 15. That fixes what a point is worth. About two-thirds of people score between 85 and 115, and a single point is a fifteenth of a standard deviation.
Against that scale the modern estimates are small. A 2015 analysis in the Proceedings of the National Academy of Sciences by Julia Rohrer, Boris Egloff and Stefan Schmukle pooled long-running household panel studies from the United States, Britain and Germany and compared siblings with each other rather than with people from other families. The firstborn advantage in measured intelligence came out at a small fraction of a standard deviation, which on the mean-100, standard-deviation-15 metric is a point or two. The same analysis found no birth-order effect on the broad personality traits — only on self-reported intellect, which is closer to the intelligence result than to a personality one.
A far larger dataset points the same way. In 2007 Petter Kristensen and Tor Bjerkedal reported in Science on the intelligence test given to Norwegian men at conscription, a record covering hundreds of thousands of individuals. Firstborns came out ahead of second-borns by something in the region of two points once the conscription score was expressed on a mean-100, standard-deviation-15 metric, with a smaller further step down to the third. Whether the true figure is nearer one point or nearer two is genuinely disputed. What recurs across datasets is the sign and the order of magnitude, not the decimal.
The dissent is real too. In 2000 Joseph Lee Rodgers and colleagues argued in American Psychologist that much of the apparent birth-order and family-size pattern was an artifact of comparing across families, and that within-family analyses of large American survey data showed no birth-order effect on intelligence worth speaking of. That paper is why this article is filed as contested rather than settled. The defensible summary is that within-family designs usually find a small firstborn advantage, sometimes find none, and never find a large one.
The firstborn advantage is real, and it is a point or two on a mean-100, standard-deviation-15 scale. Both halves of that sentence matter, and the retelling keeps only the first.
There are two ways to ask whether firstborns score higher, and they are not two versions of the same question. A between-family comparison takes every firstborn in a sample and every second-born and compares the two groups. A within-family comparison takes siblings from the same household and compares them with each other. The first is the obvious design and the wrong one.
It is wrong because the two groups differ by far more than birth order. A fourth-born child can only exist in a family with at least four children, while a firstborn exists in every family including the ones that stop at one. So a between-family comparison of firstborns against later-borns is partly a comparison of small families against large ones — and family size is not distributed at random. It tracks parental education, income, historical period and the parents' own test performance. Any of those can manufacture a gap that has nothing to do with the order the children arrived in.
That is exactly what happened to the classic evidence. In 1973 Belmont and Marolla published in Science an analysis of close to four hundred thousand Dutch nineteen-year-olds tested at conscription on Raven's Progressive Matrices, a non-verbal reasoning test built from visual patterns. Scores declined with birth order and declined with family size in a strikingly regular staircase. Two years later Robert Zajonc and Gregory Markus offered an explanation in Psychological Review: the confluence model, on which the average intellectual level of a household falls as each new baby joins it, so later children develop in a thinner environment. It was elegant, it fitted the curve, and it rested on a comparison that could not separate birth order from family size in the first place.
When the same question is asked inside families, the staircase flattens dramatically. It does not vanish in most analyses — that is the finding. But the gap between what the two designs report is larger than the effect either of them is measuring.
If the effect lives inside families, then whatever causes it has to be something that differs between siblings raised in the same house. The most informative evidence on that also comes from the Norwegian conscription records. Kristensen and Bjerkedal separated biological birth order from social rank in the family by looking at men whose older siblings had died in infancy — men who were biologically second- or third-born but who grew up occupying the position of the eldest. They scored like firstborns. Whatever is doing the work appears to be about the role a child holds in the family rather than about the pregnancy.
That narrows the candidates:
None of these has been isolated. They are not mutually exclusive, and a gap this small does not need a large cause: it could be several small ones pulling the same way, or one moderate one partly cancelled by another. Anyone who names the mechanism confidently has gone past the evidence.
Here is why none of this will ever show up in your own family. Every test score is an estimate, and test manuals publish a figure describing how much that estimate wobbles: the standard error of measurement, which is the typical distance between a person's reported score and the score they would average over many sittings. On a well-built individually administered test the standard error for a full-scale score is usually somewhere around two to three points on the mean-100, standard-deviation-15 scale. That is why manuals report a confidence band rather than a point, and why the band is commonly about ten points wide at 95 per cent confidence. Our note on what accuracy means for a test score works through the same figure from the other end.
Now set the effect against the band. A gap of one to two points is no larger than the standard error on a single administration. Two siblings tested once each could differ by that much in either direction for reasons that have nothing whatever to do with which of them was born first. In percentile terms — a percentile being simply the share of the reference sample scoring at or below a given point — a score of 100 sits at the 50th, and a score of 102 sits a little above the 55th. You can check both on our IQ percentile calculator, and see the shape they come from on the bell curve page. A five-percentile shift in a group average is a genuine population fact. It is not a fact about a person.
The overlap makes the point better than any single figure. If two groups differ by a tenth of a standard deviation, then drawing one person at random from each and asking which scores higher gives the higher-average group the win just under 53 times in 100 rather than 50. That is what a real, replicated, tiny effect looks like from the inside: almost exactly a coin toss.
It is worth being precise about how this differs from a claim that simply failed. The Mozart effect shrank towards zero as better studies arrived, and what residue remained was explained by something other than the music. Birth order did not do that. It shrank and then stopped, at a small value that keeps reappearing whenever siblings are compared with each other. The error in the popular version is not that the effect is fictional. It is that a difference of a point or two has been retold as a difference in kind.
What is worth taking from all this is the method rather than the finding. When a claim about intelligence comes with a number attached, the question that settles most of it is what was compared with what. Birth order is the cleanest available demonstration: the same children, the same test, two defensible-looking designs, and answers that differ by more than the effect itself. If you want a score of your own that you can reason about, take a properly structured test once under decent conditions rather than repeatedly under bad ones, and read the result as a range on a named scale.
The habit that survives this article is the one to keep. Before accepting any claim that something moves intelligence, ask which comparison produced it — because between families and within families give different answers about the same children, and only one of them is answering the question you asked. Then check any score you already hold on its own scale, as a percentile with a confidence range around it, rather than as a bare number.
Find your IQ score now! →So: does birth order affect IQ? On the current evidence, yes, slightly, and probably through the position a child occupies rather than the order of the births themselves. The size is one to two points on a mean-100, standard-deviation-15 scale, the exact figure is disputed, and it exists only in comparisons that hold the family constant. Every clause in that sentence has to survive the retelling for the claim to stay true, and in the version most people have heard, none of them does.
Slightly. Analyses that compare siblings within the same family find firstborns ahead on measured intelligence by roughly one to two points on a scale with a mean of 100 and a standard deviation of 15 — around a tenth of a standard deviation. The effect is replicated but its exact size is disputed, and it is far too small to explain the difference between two particular people.
On average, by a point or two, and not in any way you could observe. A one-to-two-point difference on a mean-100, standard-deviation-15 scale is no larger than the standard error of measurement on a single test administration, which is typically around two to three points. Any two siblings can differ by more than the effect for purely random reasons.
Because they compare across families rather than within them. A between-family comparison puts all firstborns against all later-borns, but later-borns only exist in larger families, and family size tracks parental education, income, era and the parents' own test scores. That design measures birth order and family background together and cannot separate them.
Much less than cross-sectional data suggests, and possibly not at all. The steep decline with family size in classic surveys largely reflects differences between the families themselves rather than an effect of siblings. Within-family analyses flatten it substantially, and an influential 2000 paper in American Psychologist argued the pattern is mostly an artifact of the comparison.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
The 1993 experiment behind three decades of "classical music makes babies smarter" ran for ten minutes, used college students, measured a single spatial task, and never mentioned IQ at all.
Almost every eleven-year-old in Scotland sat the same test on the same day in 1932. Decades later those scores turned out to track survival — and the argument about what that means has run ever since.
The numbers attached to famous names travel for decades without ever acquiring a test record, a date or a scale. Where the trail can be followed at all, it usually ends at a magazine paragraph.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →