Sitting the same test twice usually produces a higher second score. Almost none of that gain is ability — which is why employers use alternate forms, and why the second number is a worse estimate than the first.
Employers who use cognitive ability tests are measuring something that is supposed to be fairly stable. Candidates preparing for those tests are trying to raise a number. Both can be right at the same time, because the number and the thing are not the same object — and practice moves one of them a great deal more than the other.
The best evidence on retesting in real selection settings comes from a meta-analysis by John Hausknecht and colleagues, published in the Journal of Applied Psychology in 2007. Pooling studies in which candidates sat a cognitive test more than once, they found an average gain on the second sitting of roughly a quarter of a standard deviation. On a scale where the standard deviation is 15 points, a quarter of a standard deviation is about four points.
Two things moderate it consistently. Reusing the identical form produces a larger jump than moving to an alternate form built to the same specification — which is the signature of familiarity rather than improvement. And the gain is front-loaded: the step from a first sitting to a second is the big one, and each further attempt adds less.
The same shape appears on clinical IQ batteries. Test manuals publish retest data, and the gains are reliably larger on the nonverbal, speeded and puzzle-like subtests than on vocabulary, general knowledge and other stored-knowledge measures. That asymmetry is itself a clue about mechanism: what improves is whatever benefits from having seen the format before.
What is not on that list is reasoning ability. Nothing in the retest literature suggests that sitting a test repeatedly improves the general capacity the test was built to estimate.
A score raised by familiarity with one form is a less accurate estimate than the score it replaced, not a more accurate one.
A test works by comparing your performance against a reference sample who sat it under standard conditions: one sitting, no prior exposure. Arrive having already seen the format and you are no longer in the situation those norms describe. The score goes up, and its meaning as a percentile against that sample goes down.
This is why serious selection processes use alternate forms, impose waiting periods between attempts, and sometimes disregard a repeat score entirely. It is not suspicion of the candidate. It is the same logic that makes a supervised administration count and an unsupervised one not.
If the reason for practising is to remove the novelty rather than to game a number, one honest full-length run under timed conditions does most of the work. Sitting a properly built assessment once, under real conditions, tells you where you actually stand — and the result is worth reading as a percentile on a named scale with its confidence range, not as a single figure to be improved.
Find your IQ score now! →One more thing gets lost in the discussion of gains. Every score is an estimate carrying measurement error, and test manuals publish a standard error of measurement describing how far a result is expected to move between sittings for reasons that have nothing to do with learning. On a Wechsler-type scale, a responsibly reported result is a band several points wide rather than a point.
So a candidate who scores 104 and then 108 has not necessarily gained anything at all. That difference sits comfortably inside the range two sittings would produce under identical conditions. The retest literature describes an average shift across many people; it does not guarantee that any individual's second number will be higher. Some are lower.
For what a score of this kind actually predicts once you have one, our note on workplace cognitive ability testing covers the validity evidence and the 2022 reanalysis that pulled the headline figures down.
You can raise the score, modestly. Pooled employment research puts the average second-sitting gain at around a quarter of a standard deviation — roughly four points on a 15-point scale — and most of it comes from familiarity with the format, the instructions and the pacing rather than from improved reasoning.
No — usually less. The norms describe people sitting the test without prior exposure. Once you have seen the format, your performance no longer maps onto that reference sample in the same way, so the second score is a poorer estimate of your standing even though it is a higher number.
Longer intervals produce smaller practice gains, and many employers set their own waiting period, often several months to a year. Check the specific policy first — attempting a retest inside it usually means the score is not used at all.
Nonverbal, speeded and puzzle-type sections — matrices, number series, symbol tasks. Vocabulary and general-knowledge sections improve least, because there is no format trick to learn and no shortcut to knowing a word.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
Aptitude tests are among the better-studied hiring tools, and for decades the headline validity figures were quoted with more confidence than the corrections behind them deserved. A 2022 reanalysis pulled those numbers down.
A bad night does not lower intelligence. It reliably lowers the things a timed test happens to measure most closely — sustained attention, working memory and speed — while leaving vocabulary and general knowledge comparatively intact.
Recruitment screening and candidate assessment are high-risk uses under the EU AI Act, and their obligations were set to begin on 2 August 2026. A regulation in force since late July moved that date to 2 December 2027 — while leaving one workplace ban already biting.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →