No psychological test is 100% exact. Every standardised measurement of human cognition carries a margin of error and depends partly on the conditions you took it under. Rather than claim precision we cannot demonstrate, here is exactly how accuracy is defined, measured — and limited.
Excellent 4.8/51 million+ test takersIIF certified
Published by the IQ Metrics Research Team · Psychometric standards document
Any observed test score can be thought of as two things added together: your true score — the value you would get if measurement were perfect — and a quantity of measurement error. Nobody can see those two parts separately. What psychometrics can do is estimate how large the error component typically is, and that estimate is the standard error of measurement (SEM).
The formula says something intuitive: the more reliable a test is, the smaller its measurement error. A test with a reliability of 0.90 on the standard scale gives an SEM of 15 × √0.10, which is about 4.7 points. Because roughly 95% of a normal distribution sits within two standard errors, an observed score of 115 corresponds to a 95% confidence band of about 106 to 124.
The standard deviation of the IQ scale. Fixed by the scale itself, not by any individual test.
How consistently a test measures. Well-constructed cognitive tests typically report values in this region.
The resulting typical error. Published cognitive assessments commonly sit around ±3–5 points.
This is why reporting a range is more honest than reporting an integer. A result presented as “115” invites you to read a precision the measurement does not have; the same result presented as “roughly 106–124, most likely near 115” tells you what was actually measured. It also reframes small differences correctly — a gap of two or three points between two sittings is comfortably inside the noise, not evidence that anything has changed.
One caveat worth stating plainly: the ±3–5 point figure is a general property of standardised cognitive tests, derived from the reliability values such tests typically publish. It is not a measured constant of any one assessment, and we are not presenting it as a validated statistic for ours.
Most commercial test items are private. A few research item banks are not — and that difference matters for scrutiny.
The International Cognitive Ability Resource (ICAR) is a public, openly licensed pool of cognitive-ability items developed by academic researchers so that psychometric work need not depend entirely on proprietary instruments. Because the items and accompanying data are openly available, other researchers can re-analyse them, replicate findings, and publish the statistical properties of individual questions — the kind of external scrutiny that closed commercial item banks by definition cannot receive.
That open literature is also where the vocabulary of good item design gets worked out. Under item response theory, each question carries at least two parameters. Difficulty (b) describes the ability level at which a person has about a 50% chance of answering correctly — it positions the item along the scale. Discrimination (a) describes how sharply the item separates people just above that point from people just below it; a steeper curve carries more information about where someone sits.
Figure 1. Item characteristic curves. Raising difficulty (b) slides a curve to the right, so a higher ability is needed for an even chance of success. Raising discrimination (a) steepens it, so the item distinguishes more sharply between adjacent ability levels. Curves are illustrative of the standard two-parameter model, not measured from any specific item bank.
To be explicit about what this section does and does not claim: ICAR is described here as an example of open psychometric practice, and as the setting in which these design concepts are publicly documented. Nothing above should be read as a statement that IQ Metrics has been validated by, endorsed by, or formally derived from ICAR. Any such claim would need a published study behind it, and we are not making one.
Measurement error is not only statistical. Some of it is you, on a particular afternoon, in a particular room.
Sleep deprivation and acute stress can temporarily depress performance, particularly on timed tasks that lean on concentration and working memory. The effect is on the day’s performance, not on underlying ability.
Interruptions, notifications and background noise cost attention — and attention is exactly what timed reasoning items consume. A quiet, uninterrupted sitting is the single easiest thing to control.
Verbal items are sensitive to language proficiency and culturally specific knowledge. Non-verbal, matrix-style items reduce that dependence — though, as with any test, they do not remove it entirely.
Retaking a test too soon measures familiarity, not intelligence.
Sitting the same or a very similar assessment again within a short window tends to raise the observed score — test–retest research commonly reports gains in the region of 3–6 points — because you have already seen the item formats and worked out what they are asking. Your reasoning capacity has not changed in the interval; your familiarity with the test has. Spacing attempts by one to three months lets that familiarity fade. If you want to keep working in between, untimed practice on individual reasoning domains is a better use of the time than repeating the full assessment.
Both are legitimate. They answer different questions — and only one of them is a diagnosis.
| Feature | IQ Metrics online assessment | Clinical in-person battery |
|---|---|---|
| Purpose | Self-administered cognitive screening | Comprehensive clinical evaluation |
| Administration | Online, self-paced | Administered by a licensed psychologist |
| Accessibility | Immediate access from any device | Appointment and professional fee |
| Reporting | Composite score, percentile and cognitive profile | Comprehensive professional report |
| Suitable for | Personal insight and tracking your own reasoning practice | Educational, medical or clinical decisions |
Clinical batteries such as the WAIS-IV are named here only as familiar examples of professional in-person assessment. Naming one is not a claim that this assessment has been compared or validated against it.
An online assessment can give you a reasonable estimate of general reasoning ability, along with a breakdown of where your relative strengths sit. What it cannot do is function as a professional evaluation. It is not administered by a clinician, it cannot observe how you approach a problem, and it cannot diagnose anything. If you need a cognitive assessment for an educational placement, a workplace accommodation, or any medical question, that requires a qualified professional — and no online result, ours included, is a substitute for one.
A well-constructed online assessment that applies structured psychometric principles can give you a useful estimate of your cognitive ability, reported with an appropriate margin of error. It remains a screening and self-assessment tool rather than a clinical evaluation, and it should not be used to diagnose anything or to make educational or medical decisions. Our How IQ Tests Work page covers the methodology behind that estimate.
We suggest leaving at least one to three months between full assessments. Retaking the same or a very similar test sooner tends to raise the score through familiarity with the item formats rather than through any change in underlying ability. If you want to keep working in the meantime, untimed practice on individual reasoning domains is a better use of the interval than repeating the full assessment. Our IQ Score Over Time tool is built for tracking exactly this.
Several things differ between tests. Some report on a scale with a standard deviation of 15 and others use 24, so the same relative standing produces a different number. Tests also weight reasoning domains differently, vary in how carefully they are standardised, and are taken under different personal conditions such as fatigue, stress or interruption. Comparing raw numbers across tests without checking the scale is misleading. Our IQ Score Converter handles this conversion directly.
It’s the margin of uncertainty around a single result — your true ability most likely sits within a band around the reported number, not at that exact point. A tighter band means a more precise instrument, but no test, including ours, reports a single number with zero error. Our IQ Bell Curve page shows how a score maps onto a range rather than a point.
When item statistics and scoring logic are open to scrutiny, independent reviewers can check that the difficulty calibration and standardisation actually hold up, rather than taking a publisher’s accuracy claims on faith. It’s a transparency safeguard, not a guarantee of a higher score. See the broader methodology on our IQ Test Guide.
Yes — acute stress, sleep deprivation and interruption all measurably reduce working memory and processing speed on the day, independent of your underlying reasoning ability. This is one of the main reasons a single result should be read as an estimate rather than a fixed number. Our Can You Improve Your IQ page covers what actually helps versus what doesn’t.
No. A clinical evaluation is administered and interpreted by a licensed professional, often combines several instruments, and can be used for formal diagnosis or accommodation decisions. An online screening assessment estimates general reasoning ability and is not a substitute for that process. See our Why Us page for what our assessment is built to do.
Because percentile bands compress heavily near the middle of the distribution and stretch out toward the extremes, a small shift in raw score can move your percentile by a noticeable amount in the mid-range but by very little near the tails. Our IQ Percentile Calculator lets you see this directly for any given score.
Norm samples are typically largest around common working-age brackets and thinner at the very young and very old ends, which can make scores at the age extremes carry a slightly wider margin. Our IQ and Age tool breaks down how age-based norming works.
Only after converting both to the same scale — a raw number alone doesn’t tell you whether two tests agree, since they may differ in standard deviation, item content and norm population. Our IQ Comparison tool is built specifically to make that comparison meaningful.
Childhood results generally carry more measurement uncertainty than adult ones, because cognitive abilities are still developing and a young child’s attention span and test-taking familiarity add extra variability on top of the trait being measured. Our Kids IQ Test is normed specifically for this age range rather than reusing the adult scale.
Likely yes, and this is expected rather than a sign the first result was wrong. Extreme scores are more affected by chance and off-day factors, so a retest tends to land closer to the true average — a well-documented statistical pattern called regression to the mean, not evidence of a real change in ability. See our Regression to the Mean page for how this plays out in practice.
You get a composite score, your percentile, and a breakdown by reasoning domain — reported with the limits of the measurement stated openly rather than hidden.
Take our standardised assessment →