Aptitude tests are among the better-studied hiring tools, and for decades the headline validity figures were quoted with more confidence than the corrections behind them deserved. A 2022 reanalysis pulled those numbers down.
If you have applied for a graduate scheme, a call-centre role, a trainee programme or almost anything in logistics or finance, you have probably sat a timed reasoning test with a vendor logo on it. They are ordinary infrastructure now. What they actually tell an employer is less settled than the marketing around them suggests — and less settled than it was five years ago.
Most of them measure what researchers call general mental ability, usually through some mix of verbal reasoning, numerical reasoning and abstract or inductive reasoning under time pressure. They are close relatives of the reasoning sections of an IQ test, and often share item formats with them. They are not the same product: an employment test is validated against job outcomes and is usually normed on an applicant population rather than the general public, which is why a percentile from one does not transfer to the other.
A percentile, here, means the share of the reference group scoring below you. Which group is the reference is the whole question. Scoring at the 70th percentile against a general-population sample and against a sample of people who already applied for a competitive graduate role are very different results with the same label.
For roughly twenty-five years, discussions of hiring tests leaned on a single widely cited meta-analysis reporting that general mental ability was among the strongest available predictors of job performance, with a validity coefficient around the mid-point of the scale. That figure was corrected — adjusted upward using standard statistical procedures meant to compensate for measurement error and for the fact that studies only observe people who were actually hired.
In 2022 a reanalysis published in the Journal of Applied Psychology argued those corrections had been applied too aggressively and in some cases to the wrong quantities, and that the corrected validities for most selection procedures — cognitive tests among them — should be substantially lower than the figures in circulation. The reanalysis did not claim the tests are useless. It claimed the published numbers overstated them, and that several other methods, notably structured interviews and work samples, sit closer to cognitive tests than the old table implied.
This is a live disagreement among specialists rather than a settled overturning, and it is worth stating that plainly. What is safe to say is that anybody quoting a single decisive validity figure today — in either direction — is ahead of the evidence.
Even on the most favourable published figures, a test score left most of the difference in job performance between two people unexplained.
Employment testing in the US sits inside a specific regulatory structure. The Uniform Guidelines on Employee Selection Procedures, adopted in 1978 by the federal enforcement agencies, set out how a selection procedure that produces a substantially different pass rate across race, sex or ethnic groups must be justified — essentially, by evidence that it is job-related and consistent with business necessity. The doctrine behind them comes from Griggs v. Duke Power Co. (1971), in which the Supreme Court held that a practice neutral on its face can still be unlawful if it screens out protected groups and cannot be shown to relate to job performance.
None of this makes cognitive testing unlawful, and it is used lawfully at scale. It does mean that the burden of showing a test measures something the job needs sits with the employer, and that “it correlates with performance in general” is not the same as “it is defensible for this role”.
An employment aptitude test and a general IQ assessment are close cousins with different reference groups, so a percentile from one does not carry over to the other. If you want your own number on the general-population scale rather than an applicant pool, that is a different measurement: a score on a named scale, its percentile, and the confidence range around it.
Find your IQ score now! →The reasonable position, given where the evidence sits, is that these tests carry real information and less of it than the marketing claims — and that the honest use of one is as a single input alongside structured interviews and work samples, not as a ranking of candidates by a number.
They overlap heavily in item format and both load on general reasoning, but they are validated and normed differently. An employment test is validated against job outcomes and usually normed on applicants, so its percentiles describe your standing among other candidates rather than among the general population.
Yes, for that test. Familiarity with the format and the timing reliably improves performance on that item type, which is a real advantage in a hiring process. It is not the same as a change in general reasoning ability, and honest practice providers say so.
In the United States, yes, within limits. Under the Uniform Guidelines and the case law from Griggs v. Duke Power, a selection procedure that produces a substantially different pass rate across protected groups has to be shown to be job-related and consistent with business necessity. Rules differ by country.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
On the scale most modern tests use, 120 sits around the 91st percentile. Change the scale and the same number moves. Add the measurement error every test carries and it stops being a point at all.
Two people can perform identically on two different online tests and walk away with scores twenty points apart. The scale, the comparison group and the margin of error are what separate an assessment from a quiz.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →