IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

Test Anxiety and IQ Scores

Mind & Everyday Life

How Many IQ Points Does Test Anxiety Actually Cost?

Almost everyone who has had a disappointing result has wondered whether nerves cost them fifteen or twenty points. The evidence says no. Test anxiety and cognitive performance correlate at roughly r = -0.20, which works out to a few points on a 15-point scale. Here is the number, and what it changes.

Bar chart comparing the twenty-point loss people claim test anxiety costs with the roughly three points the measured effect size implies

Can anxiety lower your IQ score? Yes — but not by anything close to the margin most people assume. Studies that measure test anxiety alongside cognitive test performance find a relationship of roughly r = -0.20, which translates to something on the order of three points on the familiar 15-point scale. That is a few points between an average test-taker and a highly anxious one, not fifteen and not twenty. The rest of this article shows where that figure comes from, why the effect is small but genuinely real, and what it should change about how you read your own result.

Where the fifteen-point claim comes from

Nearly everyone who has sat a supervised test has said some version of it: I would have scored fifteen or twenty points higher if I had not been so nervous. The thought is appealing because it explains away a disappointing number without asking anything of the person who got it.

The problem is scale. Fifteen points is a full standard deviation — the distance between the middle of the distribution and the edge of roughly the top sixteen percent of test-takers. Ordinary pre-test nerves do not move a person that far. Nothing in the measured relationship between test anxiety and test performance is anywhere near large enough to produce a shift of that size.

What the evidence does support is a smaller, stubborn, fairly consistent drag. Small effects are still real effects. They are just not the ones that rewrite a life story.

What the research actually finds

Meta-analyses pooling many studies of test anxiety and performance on cognitive and academic tests tend to land on a correlation in the region of r = -0.20. Much of that literature is built on school and university assessments rather than on standardised intelligence tests, so carrying the figure across to IQ points is a reasonable estimate rather than a direct measurement. Estimates range roughly from -0.15 to -0.25, and much of that spread is not noise: it moves with who was sampled, how anxiety was measured, and how much the test genuinely mattered to the person sitting it.

Bar chart comparing the twenty-point loss people claim test anxiety costs with the roughly three points the measured effect size implies
Bar chart comparing the twenty-point loss people claim test anxiety costs with the roughly three points the measured effect size implies

Three things push the estimate around, and they are worth knowing before you take any single figure too seriously:

  • How anxiety was measured. Questionnaires filled in after the test pick up disappointment as well as anxiety. Physiological measures of arousal track performance more weakly than self-reported worry does, so the method chosen moves the estimate before any real difference does.
  • How high the stakes were. A university entrance exam tends to generate more anxiety, and a larger measured effect, than a practice test taken at home out of curiosity.
  • Who was in the sample. Student samples dominate this literature, and students are not a random slice of the population, which limits how far any one estimate travels.

A correlation of -0.20 means anxiety is associated with something like four percent of the variation in scores across a group. The other ninety-six percent of the differences between people sits somewhere else entirely. That is the frame to hold before anyone converts anything into points.

Turning that correlation into IQ points

Correlations are hard to feel; points are easy. So here is the translation, with the reasoning shown, precisely so you can see how approximate it is.

Most IQ scales are built with a standard deviation of 15 points. If anxiety and score are related at about -0.20, a person one standard deviation above average in test anxiety is expected to score roughly 0.20 × 15, or about 3 points, below a person of average anxiety. Nothing in that figure is held constant: it is a raw association, not an isolated effect of anxiety. Push it to two standard deviations above the mean in test anxiety, which is uncommon, and the expected gap is around six points — though the linear assumption behind that extrapolation is doing a lot of work out at the tail.

So the honest headline is a few points, not twenty. Every word there is doing work. It is an average across groups, not a prediction about any one person, and it applies to people whose anxiety slows them down rather than stopping them. For the minority who freeze or leave a test unfinished, the loss can be far larger, because what they end up with is not a measure of their ability at all.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

Why the cost is real but small

The best-supported explanation is crowding. Working memory — the small, temporary workspace where you hold a matrix pattern, a half-finished sequence, or the three candidate answers you are weighing against each other — has a hard capacity limit. Worry is not silent: it occupies that same workspace with self-monitoring, clock-watching and a running commentary on how badly things are going.

That account makes a prediction that can be checked, and it broadly holds up. The cost concentrates on items that load working memory heavily and on tests run against a clock, which is one reason a timed IQ test is likelier to expose anxiety than an untimed one. On easy items, or when you can take as long as you like, there is spare capacity to absorb the interference.

This is the same mechanism driven by a different stressor as the one behind sleep loss and test performance. Fatigue and worry both draw down the same limited resource, and both produce modest average losses concentrated in the same kinds of items. Read that piece as the companion to this one.

There is a trait side to all of this too — some people are dispositionally more anxious, and that disposition travels with a cluster of habits around testing. Rather than re-derive it here, see how personality relates to measured ability.

Anxious test-takers also behave differently

The raw observed gap between anxious and non-anxious test-takers is not purely anxiety pressing on cognition in the moment. Anxious people also do different things, and those things land in the score just as surely.

  • They prepare differently. Avoidance is a normal anxiety response, so some anxious test-takers do less familiarisation rather than more, and arrive facing a format they have never seen.
  • They start slower. Time spent re-reading the instructions and second-guessing item one is time not spent on items twelve through twenty.
  • They abandon hard items sooner. Giving up early is a behaviour, not a capacity limit — but on a scored test it looks identical to being unable to solve the item.
  • They are likelier to quit part-way. A test abandoned halfway does not measure ability. It measures how long the person stayed.

Each of those routes lowers a score without anxiety having interfered with reasoning directly. Which means some share of that already-modest -0.20 is behaviour rather than anxiety acting on reasoning in the moment, so the direct cognitive cost is probably smaller still. How much smaller is not something a correlation on its own can tell you.

That is better news than it sounds. Behaviour is a great deal easier to change than temperament, and the routes above are the ones a second, better-organised sitting can actually close.

What a few points should and should not change

An expected loss of around three points is a reason to consider retesting. It is not a reason to discard the result you have. If your score landed near a boundary you care about and the sitting was a genuinely rattled one, that is a legitimate argument for a second attempt.

It is worth being specific about what near a boundary means. A three-point expectation matters if your result sits a handful of points from a threshold you are using for something; it matters very little if you are twenty points away, because a calmer sitting is not going to carry you across a gap that size.

Go in knowing what a retest actually buys. Scores tend to rise on repeat testing for reasons that have nothing to do with your nerves improving: the practice effect on cognitive ability tests is well documented, so a higher second score is partly familiarity with the format and, if the first sitting was unusually low for you, partly ordinary regression toward your own average — and only partly a calmer state of mind.

It also helps to know how much any single sitting can move for perfectly mundane reasons. The guide to IQ test accuracy covers the measurement error surrounding one score. Anxiety is one contributor sitting inside a band that is wider than most people expect, and this article is only quantifying that one contributor.

What actually helps before a test

The interventions with the strongest support are unglamorous: familiarise yourself with the format, and practise under conditions that are genuinely timed. Both work for the same reason — they strip out the novelty and the surprise of time pressure that generate the anxiety in the first place, rather than trying to manage the feeling once it has already arrived.

In practice that means a full-length run with the clock genuinely running rather than a relaxed browse through sample items. The preparation guide and the walkthrough of how a test is structured cover those mechanics, so this piece will not repeat them.

None of this makes anxiety trivial. It makes it tractable. If you suspect nerves cost you something last time, the useful next step is a second run under conditions you control, with the clock running honestly. Take the IQ test that way and compare the two results, holding in mind that a few points of movement in either direction is exactly what the evidence predicts.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged IQ Score, IQ Test, iq test results, iq test score, iq testing, Measurement Error, mental ability, personal development, Practice Effect, Processing Speed, test anxiety, Test Preparation, Working Memory

IQ and Creativity

Research & Evidence

Does Intelligence Stop Predicting Creativity Above IQ 120?

The threshold hypothesis says IQ and creativity move together up to about 120 and barely at all above it. It is a real claim with a real number, and modern re-analyses have weakened it. Here is what the segmented-regression work found, why the two kinds of task differ, and why measurement noise matters.

Chart of creative achievement against IQ, contrasting the threshold claim that the line goes flat above 120 with re-analyses showing it keeps rising more gently

It fades, but it does not stop. IQ and creativity correlate positively but modestly across ordinary score ranges, and the relationship may well weaken somewhere in the 110 to 130 band. What it does not appear to do, on the best modern evidence, is switch off: the correlation shrinks rather than vanishing, and where it starts shrinking moves with the creativity measure you use.

Two things make the question stubborn. Creative tasks ask the mind for a different operation than a reasoning item does, and the instruments that score them are far noisier than an IQ test.

The threshold hypothesis, stated as a number

The claim is most often associated with Ellis Paul Torrance, though J. P. Guilford and others in the same 1960s literature advanced versions of it too. Stated plainly: intelligence is necessary but not sufficient for creative performance, so below roughly IQ 120 more measured ability tends to mean more creative output, and above that point extra points buy little or nothing.

There is nothing magical about 120 itself. On a scale with a mean of 100 and a standard deviation of 15, a score of 120 sits just above the 90th percentile, so about nine people in a hundred score higher. Where it falls on the IQ bell curve makes the point better than a sentence can, and the percentile calculator will place any other score beside it.

So the threshold, if it exists, sits at the edge of the top decile, not out among the rare scores. That matters: a claim about the top ten percent is testable in ordinary samples; a claim about the top tenth of one percent is not. Our piece on what a score of 120 actually means covers the everyday side.

  • Below the line: a clearly positive correlation between measured intelligence and scored creativity, steep enough to see in a scatter plot. For reference, meta-analytic summaries put the correlation across the whole score range at roughly 0.2, with individual estimates scattered widely on either side of it.
  • Above the line: a correlation close to zero, so that two people scoring 125 and 145 ought to be indistinguishable on creative measures.
  • A visible kink: a scatter plot whose slope changes at a particular score, not a straight line that merely drifts.

What happened when the claim was re-tested

The threshold hypothesis is unusually easy to test, because it makes a geometric prediction. Fit one line to the data and it should fit badly. Fit two lines joined at a breakpoint estimated from the data itself — segmented, or piecewise, regression — and the second line should come out flat. Several groups have done exactly that, on Torrance-style divergent-thinking batteries and on fresh community samples.

Chart of creative achievement against IQ, contrasting the threshold claim that the line goes flat above 120 with re-analyses showing it keeps rising more gently
Chart of creative achievement against IQ, contrasting the threshold claim that the line goes flat above 120 with re-analyses showing it keeps rising more gently

The results have been mixed. Some analyses do recover a breakpoint, but its location moves with the scoring rule rather than sitting at 120: estimates have landed in the mid-80s for simple idea fluency, near 100 for originality scored across all responses, and close to 120 only when originality was judged on a person’s best ideas alone. The confidence intervals around those estimates are usually wide.

Other analyses find that a single straight line describes the data about as well as two do, which is precisely what the hypothesis says should not happen. A bend only some analysts can find, at a score that moves with the marking scheme, is a weak bend.

Longitudinal work on people far above the threshold cuts against it too. In samples selected in adolescence for very high mathematical or verbal ability, differences within that already-elite group still predicted patents, publications and creative accomplishment decades later. If ability stopped mattering above 120, those differences should have washed out. They did not.

Something does visibly change in the upper range, though, and the competing explanations are unglamorous.

  • Range restriction. Correlations shrink automatically when you slice the bottom off a distribution, whether or not the underlying relationship changed.
  • Ceiling effects. Many creativity tasks are short and easily maxed out, so strong performers pile up at the top and the spread a correlation feeds on disappears.
  • Sampling. Gifted-programme and selective-university cohorts supply most above-threshold data, and they are selected on close to the very variable under study.

Two operations: divergent and convergent thinking

The argument also refuses to settle because creativity tests do not all ask for the same thing. Two families of task dominate the literature, they behave differently against IQ, and the quickest way to see why is to imagine sitting them.

The alternate uses task

You are handed a common object — a brick, a paperclip, a newspaper — and asked to list as many uses for it as you can in a few minutes. Responses are scored on fluency (how many), flexibility (how many categories), originality (how rare the answer is against a reference sample) and elaboration. A doorstop counts. So does grinding the brick down for pigment. There is no key at the back of the book.

The remote associates task

Here you get three words — cottage, Swiss, cake — and have to find the fourth that links all three. The answer is cheese, and once you see it, it is plainly right. That single defensible answer makes the task behave much more like a reasoning item, and its correlations with general intelligence tend to run higher than those of alternate-uses scores. The word creativity is covering two rather different measurements.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

How that compares with what an IQ test asks

A matrix-reasoning item shows a three-by-three grid of shapes with the bottom right cell missing and six or eight candidates underneath; exactly one completes the rule governing the rows and columns. A number-series item gives 2, 6, 12, 20 and wants 30, because the gaps grow by two each time. In both cases the correct response can be argued from the stimulus alone, which is what makes it scorable; the mechanics are covered in how IQ tests work.

Set that beside the brick. Two competent scorers marking matrix items agree on essentially every response; two scorers marking originality will not, unless both consult the same frequency table from the same reference population. That disagreement is not sloppiness. It is built into what is being measured.

Ability, output, and the ingredients that are not cognitive

A test measures what somebody can produce in a few minutes under instruction. A creative career measures what they finished, showed to other people, defended against criticism and kept doing after the first rejection. Related quantities, certainly — but nobody should expect them to track each other tightly.

Among the non-cognitive ingredients, openness to experience is the trait most consistently linked with creative achievement, in many studies at least as strongly as measured ability is. Persistence and sheer accumulated hours inside a domain do comparable work. Our comparison of IQ and personality sets them side by side.

The broader question — what else has a claim on the word intelligence — is treated in is intelligence limited to IQ. This page is narrower and more checkable: whether one number stops tracking another above a particular value.

The reliability problem, which is the strongest argument here

Any correlation between two measures is capped by how reliably each is measured. A well-constructed supervised IQ test typically reports test-retest reliability at or above 0.90, so somebody who sits it twice lands in close to the same place. Scored creativity batteries do considerably worse, and much less consistently: depending on the task, the scoring scheme and the interval between sittings, reported retest figures run from roughly 0.5 to roughly 0.8, and no single value describes them.

The consequence is arithmetic rather than philosophical. When one of your two variables is noisy, the observed correlation is dragged toward zero even where the true relationship is strong. Part of the modest link reported between intelligence and creativity is therefore a fact about instruments rather than about minds. Statisticians call the adjustment correcting for attenuation; applying it raises the estimates, though it produces a projection of what a perfect instrument would have found rather than a fresh measurement.

This cuts in more than one direction. Noise on its own would blur the relationship everywhere rather than bend it at one particular score, so unreliability alone is not the whole story. It is not a defence of the threshold either: a ceiling on the creativity measure, of the kind described above, produces something that looks very like a kink.

What poor reliability does guarantee is that no relationship above 120 and a real relationship too faint for a blunt instrument to detect look identical in a scatter plot. Our page on IQ test accuracy sets out the error bars a single score carries.

So what does a score above 120 say about creative potential?

Less than the number’s precision implies, and more than nothing. The defensible reading is that intelligence behaves like a resource with diminishing returns for creative work: useful throughout, most decisive at the lower end of the range, and progressively outweighed further up by interest, temperament, opportunity and hours on task.

The threshold hypothesis survives as a rough description rather than a law with a fixed value. Estimates vary, breakpoints move with the measure, and several careful analyses find no clean elbow anywhere in the data. Treat 120 as a landmark in a conversation, not as a gate.

If a result put you near that band, the useful next step is unhurried: sit a properly timed IQ test under decent conditions and read the score with its error bar attached. It will tell you something real about reasoning and pattern work. What it cannot tell you is where you land relative to its own slope once temperament, interest and hours in a domain have had their say.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged Convergent Thinking, Creativity, Divergent Thinking, genius iq level, intelligence test, iq level, IQ Science, IQ Score, Measurement Error, mental ability, mental potential, power of mind, Threshold Hypothesis

IQ vs EQ

Mind & Everyday Life

IQ vs EQ: What Each One Measures and Why You Need Both

IQ measures reasoning; EQ measures how well you read and manage emotion. They are different constructs, measured in different ways, with different amounts of evidence behind them. Here is what each one actually predicts, which is easier to change, and where the popular claim that EQ matters more comes from.

Side by side comparison of what IQ measures and what EQ measures

“IQ gets you hired, EQ gets you promoted” is one of the most repeated lines in workplace writing, and like most memorable lines it is partly true and partly a simplification that hides the interesting detail. IQ and EQ are not two versions of the same thing, and they are not rivals. They measure different capacities, they are measured by very different methods, and the quality of evidence behind them is not equal.

This article sets out what each one is, how each is assessed, what each actually predicts, and which of the two you can realistically change.

What IQ measures

IQ, or intelligence quotient, is a standardised measure of reasoning ability. A modern test samples several distinct capacities — verbal comprehension, perceptual or fluid reasoning, working memory and processing speed — and combines them into a single composite score scaled so that the population average is 100.

The score is not a count of correct answers. It is a position: a statement of how far your performance sits from the average of a norm group of people your own age. That is why the same raw performance yields different numbers on different scales, and why the IQ bell curve is the right mental picture for what a score means. If the concept is new, our explainer on what an intelligence quotient is covers the foundations.

What IQ deliberately does not measure is knowledge. A good test avoids anything you could revise for, because the target is how well you handle a problem you have never encountered, not how much you have accumulated.

What EQ measures

EQ, or emotional quotient, refers to emotional intelligence: the ability to perceive emotion accurately, to use emotion to assist thinking, to understand how emotions develop and combine, and to regulate emotion in yourself and others.

The concept was introduced in the academic literature by Peter Salovey and John Mayer in 1990 and reached a mass audience through Daniel Goleman’s 1995 book. Those two origins matter, because they produced two rather different ideas that share a name — a tightly defined ability model on one side and a much broader popular bundle of traits and habits on the other.

The four-branch model

  • Perceiving emotion — reading emotional signals in faces, voices and situations, including your own.
  • Using emotion — harnessing mood to support reasoning, judgement and creativity.
  • Understanding emotion — knowing how emotions escalate, blend and change over time, and what causes them.
  • Managing emotion — regulating your own responses and influencing those of other people.

Read as a list, these look like personality. Read as abilities, they are things you can be measurably better or worse at, which is what makes the ability model testable in a way the popular version is not.

IQ vs EQ at a glance

Side by side comparison of what IQ measures and what EQ measures
Side by side comparison of what IQ measures and what EQ measures

The two constructs differ on almost every dimension that matters: what they sample, how stable they are, how they are scored, and how confident we can be in the measurement.

  • What is sampled — IQ samples reasoning under time pressure; EQ samples emotional perception, understanding and regulation.
  • Stability — IQ is highly stable across adulthood; EQ is generally more responsive to experience and deliberate practice.
  • Measurement quality — IQ tests have close to a century of psychometric development behind them; EQ measurement is younger and much less settled.
  • Scoring — IQ items have objectively correct answers; many EQ instruments do not, which is the central difficulty in the field.

How each one is measured

Measuring IQ

Individually administered tests such as the Wechsler scales and the Stanford-Binet are the reference standard. They are delivered one-to-one by a trained administrator, take one to two hours, and report a composite score alongside index scores and a confidence interval. The methodology is mature and the reliability figures are high and publicly documented. Our guide to the professional IQ tests compares the major instruments.

Measuring EQ, which is a harder problem

EQ instruments split into two families, and the split is the reason the field is contested.

Ability tests, such as the MSCEIT, present emotional problems with scored answers — identify the emotion in this face, judge which action would best regulate this feeling. They behave like tests. The difficulty is deciding what counts as the correct answer, which is usually settled by consensus or expert judgement rather than by fact.

Self-report questionnaires, such as the EQ-i and the many quizzes derived from it, ask you to rate your own emotional skills. They are quick and popular, and they have an obvious structural weakness: someone with poor emotional perception is precisely the person least equipped to rate their own emotional perception accurately. Self-report EQ scores also correlate substantially with ordinary personality traits, which raises the question of how much new ground they cover.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

Which one predicts success better?

This is where the popular claim needs unpacking. Across the research literature, measured cognitive ability is one of the more consistent predictors of academic attainment and of performance in cognitively demanding work, and its predictive power tends to increase with job complexity.

Emotional intelligence shows its clearest value in roles built around interpersonal demand — leadership, negotiation, care work, teaching, client-facing and team-based work — and in outcomes such as team cohesion and conflict handling that a reasoning test was never designed to touch.

The honest summary is that they predict different things, and the sensible question is not which is larger but which is relevant to the situation in front of you. Neither explains most of the variance in anyone’s life; motivation, opportunity, health, circumstance and luck do enormous work that neither construct captures. We explored that wider point in is intelligence limited to IQ.

Can you raise either one?

This is the most practically useful difference between them.

IQ is comparatively resistant to change in adulthood. Scores improve modestly with test familiarity, and they can be depressed by poor sleep, illness, stress or anxiety — which means removing those things can raise a measured score without changing the underlying ability at all. Our article on the factors that affect IQ test results covers what genuinely moves a number.

Emotional skills are more trainable. Perceiving emotion accurately, pausing before reacting, naming what you feel, and reading a room are all capacities that improve with deliberate attention and feedback. Whether that improvement shows up on an EQ test is a separate question from whether it shows up in your relationships, and the second one is the one worth optimising for.

Where "EQ matters more than IQ" comes from

The claim traces largely to popular writing from the mid-1990s, where figures attributing the large majority of workplace success to emotional intelligence were widely quoted. Those figures were considerably stronger than the peer-reviewed evidence supported at the time, and stronger than it supports now.

There is a second, more subtle reason the claim feels true. In selective environments — a competitive university course, a demanding profession — everyone present has already cleared a cognitive bar. Within that narrowed range, differences in reasoning ability explain less of what separates people, and interpersonal skill explains relatively more. The observation is real; the generalisation from it to the whole population is not.

How to use both

Treat them as answering different questions. If you want to know how you handle novel reasoning problems under time pressure, that is an IQ question, and it deserves a properly normed test that reports its scale, your percentile and a confidence range rather than a bare number. If you want to know why a conversation went badly or why a team is not functioning, no IQ score will help you.

If you are curious about the first, our IQ test reports all three figures, and the percentile calculator will show you what any score means as a rank. Just do not expect either to tell you how good you are at reading a room — that was never what they were built to measure.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged emotional intelligence, emotional quotient, intelligence test, IQ, iq level, IQ Score, iq vs eq, mental ability, mental capabilities, mental potential, personal development, power of mind

Mensa IQ Score Requirements

Scores & Scales

What IQ Score Do You Need for Mensa?

Mensa admits the top 2% of scorers, but "the top 2%" is a rank, not a number, and the number it maps to changes with the test you sat. Here is what qualifies on the Wechsler, Stanford-Binet and Cattell scales, the two routes to membership, and what the supervised admission test is actually like.

Bell curve showing the top 2 percent of IQ scores that qualify for Mensa

Mensa is the oldest and best known high-IQ society in the world, and the question people ask about it more than any other is the simplest one: what score do you actually need? The short answer is the 98th percentile — the top 2% of the population. The longer answer is the one that helps, because a percentile is a rank, not a score, and the number that rank corresponds to changes depending on which test you sat.

That is why you will see 130, 132 and 148 all quoted as “the Mensa score” by sources that are each telling the truth. This article explains the one requirement that never changes, why three different numbers can describe the same performance, and how the two routes into Mensa differ.

The one requirement that never changes

Mensa has exactly one entry criterion: a score at or above the 98th percentile on an approved, supervised test of intelligence. Nothing else counts. Your age, education, occupation, nationality and income are all irrelevant to eligibility, and there is no interview, no essay and no recommendation.

Two words in that criterion do most of the work. Supervised rules out anything you take unproctored at home, including every free online quiz. Approved means the test appears on the list your national Mensa maintains — that list is genuinely long, running to dozens of instruments, but it is finite and it is published.

The 98th percentile itself is worth pausing on. It means that out of a hundred people drawn at random from the general population, roughly ninety-eight would score below you. It does not mean you answered 98% of the questions correctly, which is the single most common misreading of any percentile figure. If you want to see how a given score translates into a rank, the IQ percentile calculator does the conversion directly.

Why the qualifying number is not always 130

Bell curve showing the top 2 percent of IQ scores that qualify for Mensa
Bell curve showing the top 2 percent of IQ scores that qualify for Mensa

An IQ score is not a count of correct answers. It is a statement about how far your performance sits from the average, measured in standard deviations. The average is fixed at 100 by convention on essentially every modern test. The standard deviation is not fixed at all: different publishers picked different values decades ago and never converged.

The 98th percentile sits about 2.05 standard deviations above the mean. Multiply that distance by whatever standard deviation your test uses, add 100, and you get the qualifying score for that test.

Tests scaled to a standard deviation of 15

The Wechsler family — the WAIS for adults and the WISC for children — uses a standard deviation of 15, and so do most tests published in the last forty years. On this scale the 98th percentile lands just under 131, and the qualifying score is commonly listed as 130.

Tests scaled to a standard deviation of 16

The Stanford-Binet Intelligence Scales use a standard deviation of 16. The same performance that produces 130 on a Wechsler test produces roughly 132 here. Nothing about the person changed; only the ruler did.

The Cattell III B and its standard deviation of 24

British Mensa’s own supervised test reports on the Cattell III B scale, which uses a much wider standard deviation of 24. The qualifying score there is 148. This is the single biggest source of confusion about Mensa scores, and it is why someone quoting “148” is not exaggerating — they are quoting a different scale.

The same person, three different numbers

It is worth stating the consequence plainly. A person who performs at exactly the 98th percentile would be reported as roughly 130 on a Wechsler test, 132 on a Stanford-Binet, and 148 on the Cattell III B. All three numbers describe one identical level of performance.

This is also why comparing your score against a friend’s is meaningless unless you both know which scale each number came from. A score quoted without its standard deviation is an incomplete piece of information. The IQ score converter exists for exactly this problem: give it a score and the scale it was measured on, and it will tell you what the equivalent is elsewhere.

  • A score of 130 on an SD-15 test is not below a score of 132 on an SD-16 test.
  • A score of 148 on the Cattell scale is not near-genius on a Wechsler scale.
  • Any score without a named scale cannot be compared with any other score.

The two routes into Mensa

Every national Mensa offers the same two paths, and you only need one of them.

Route one: sit Mensa’s supervised admission test

This is the standard route and the one most applicants use. You book a session, attend in person at a scheduled testing location, and sit a proctored battery under timed conditions. There is a modest administration fee, and the result comes back as a straightforward qualified or not qualified, usually within a few weeks.

Most national chapters allow only a limited number of attempts in a lifetime — often just one or two — precisely because practice effects would otherwise inflate scores on repeat sittings. Check your own chapter’s rule before you book, because it is not the kind of thing you can undo.

Route two: submit evidence from a test you have already taken

If you have previously been assessed on an approved test — often through a school psychologist, an educational assessment, a clinical evaluation or an occupational selection process — you can submit that documentation instead of sitting anything new. Mensa reviews the paperwork and admits you if the score clears the threshold.

The evidence has to be the original scored report from a qualified administrator, not a self-reported number, and some chapters will not accept results from before a given date because the test norms have since been restandardised. That restandardisation is itself a real phenomenon worth understanding, and we cover it in our article on the Flynn effect.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

What the supervised admission test is like

The details vary by country, but the shape is consistent. You should expect a battery lasting somewhere between forty-five minutes and two hours, delivered under strict time limits, with sections that lean heavily on pattern recognition, sequence completion, verbal analogies and logical reasoning.

Two features surprise people. The first is the pace: the tests are deliberately built so that almost nobody finishes every item, and being unable to complete a section is normal rather than a sign of failure. The second is the absence of general knowledge. You are not being asked what you have learned; you are being asked how quickly you can work out something you have never seen before. If you want a sense of the item formats in advance, our guide to the types of questions asked on an IQ test walks through the common categories.

Can you prepare for it?

Partly, and it is worth being precise about which part. You cannot meaningfully raise the underlying reasoning ability the test is trying to measure in a few weeks of study. What you can do is remove the friction that costs people points for reasons that have nothing to do with ability.

  • Learn the item formats, so you are not decoding instructions while the clock runs.
  • Practise working at speed, because unfamiliar time pressure is itself a handicap.
  • Sleep properly the night before — fatigue is one of the few factors with a reliably measurable effect on test performance.
  • Sit the test at a time of day when you are normally alert, not squeezed around something stressful.

Our article on how to prepare for an IQ test goes into this in more depth, including the factors that genuinely move a score and the ones that do not.

What Mensa membership is, and what it is not

Mensa is a social and intellectual society. Membership gives you access to local and national groups, special interest groups covering an enormous range of subjects, publications, and international gatherings. Many members join for the community rather than the credential.

It is not a professional qualification, it is not recognised by employers as one, and it does not certify competence in anything. It certifies a single test performance on a single day. That is worth remembering both if you qualify and if you do not, because a great deal of what people actually want from a high score — confidence, credibility, opportunity — is not what the score is for. We looked at this gap more broadly in is intelligence limited to IQ.

If you fall short of the threshold

Roughly 98 people in every 100 do, which is the entire point of the criterion. A score below the cutoff is not a verdict on your intelligence, your capability or your prospects; it is a statement about where one timed performance sat relative to a norm group.

It is also worth knowing that scores carry a margin of error. A well-constructed test typically reports a confidence interval of several points either side of the number, so someone who scores 127 and someone who scores 131 are not reliably different people — they are two draws from overlapping ranges. Any score reported without that range is being presented more precisely than it deserves.

Checking where your own score lands

If you want to know your position before committing to a supervised sitting, the sensible first step is a properly normed assessment that reports the three figures that make a score interpretable: the scale it was measured on, the percentile it corresponds to, and the confidence range around it. Our IQ test reports all three, and the IQ tools collection lets you convert between scales and see where any score sits on the distribution.

That will not qualify you for Mensa — nothing unsupervised can — but it will tell you whether a supervised sitting is worth the fee, which is the practical question most people are really asking.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged genius iq level, high iq society, high iq test, iq classifications, iq highest score, iq level, iq rank, iq rating, IQ Score, IQ Test, IQ Test Range, iq test ranges, iq test score, mensa, mensa iq score, standard deviation iq

Professional IQ Tests

Taking a Test

Professional IQ Tests: What Psychologists Actually Use

When a psychologist measures intelligence, they reach for one of a small number of standardised instruments. This guide covers what the Wechsler scales, Stanford-Binet, Woodcock-Johnson and Raven’s Progressive Matrices each measure, who each one is for, and where an online test genuinely fits.

Comparison of the professional IQ tests psychologists use and what each measures

“IQ test” is a category, not a product. When a psychologist actually measures intelligence, they reach for one of a small number of standardised instruments, and those instruments differ in how they are delivered, what abilities they sample, how long they take, what age group they were normed on, and how heavily they lean on language. A score means very little until you know which of them produced it.

This guide walks through the tests professionals actually use, what each is genuinely good at, and how to work out which kind of result you are looking at. If you want to compare the shorter formats available on this site instead, the IQ tests hub covers those.

The first division: individual or group

The most important split is not verbal versus nonverbal — it is whether a trained administrator sits with you.

Individually administered tests are delivered one-to-one over roughly one to two hours. The administrator controls timing, watches how you approach each item, and can tell the difference between not knowing an answer and misunderstanding the question. These are the tests used clinically, educationally and diagnostically, and they are the reference standard for accuracy.

Group tests are delivered to many people at once, usually on paper or screen, with fixed instructions and no individual observation. They are far cheaper and faster, which is why they dominate schools, military selection and large-scale screening. They are also less precise for any single individual, because none of the observational detail survives.

The major individually administered tests

Comparison of the main types of IQ tests and what each one measures
Comparison of the main types of IQ tests and what each one measures

The Wechsler scales: WAIS and WISC

The Wechsler tests are the most widely used individual intelligence tests in the world. The WAIS covers adults and older adolescents; the WISC covers school-age children; the WPPSI covers preschoolers. They are periodically revised, and each revision is renormed on a fresh standardisation sample.

Rather than producing one number, a Wechsler test reports a full-scale score built from index scores covering verbal comprehension, visual-spatial and fluid reasoning, working memory and processing speed. That structure is the real value: two people can reach the same composite by very different routes, and the profile of indexes is often more informative than the headline figure. Wechsler tests use a standard deviation of 15.

The Stanford-Binet Intelligence Scales

The Stanford-Binet is the direct descendant of the original Binet-Simon scale and the oldest continuously developed intelligence test still in use. Its current edition covers an unusually wide age span in a single instrument and assesses five factors in both verbal and nonverbal form, which makes it useful at the extremes of the range where other tests run out of items. It uses a standard deviation of 16, which is why Stanford-Binet scores sit slightly higher than Wechsler scores for identical performance — a difference our IQ score converter handles directly. The test’s long history is covered in our piece on the history of IQ testing.

The Woodcock-Johnson Tests of Cognitive Abilities

The Woodcock-Johnson is built explicitly on the Cattell-Horn-Carroll model of intelligence, which organises cognitive ability into a hierarchy of broad and narrow factors. Its distinguishing feature is that it pairs with a co-normed achievement battery, so cognitive ability and academic attainment can be compared on the same standardisation sample. That makes it a common choice in educational assessment, particularly where a learning difficulty is in question.

The Kaufman batteries

The KABC-II was designed with a specific concern in mind: reducing the influence of acquired knowledge and language on the resulting score. It offers alternative theoretical scoring models and a version that minimises verbal demands, which makes it useful for children from varied linguistic and cultural backgrounds.

Nonverbal and culture-reduced tests

Every verbal test carries an unavoidable problem: it measures language proficiency alongside reasoning. For someone tested in a second language, or from a background unlike the norm group, that confound can be large. Nonverbal tests exist to reduce it.

Raven’s Progressive Matrices

Raven’s is the best known nonverbal reasoning test and one of the purest measures of fluid intelligence available. Every item is a visual pattern with a piece missing; you choose the piece that completes the rule. There are no words, no cultural references and no general knowledge anywhere in it. It comes in progressive difficulty levels, from Coloured for young children through Standard to Advanced for high-ability adults.

The Cattell Culture Fair Intelligence Test

The CFIT was built for the same purpose using series completion, classification, matrix and conditional reasoning items, all nonverbal. It reports on a scale with a standard deviation of 24, which is why Cattell scores look dramatically higher than Wechsler ones — the ruler is wider, not the person. This is the scale British Mensa uses, as covered in our article on Mensa IQ score requirements.

TONI, UNIT and Leiter

These go further still, minimising or removing language from the instructions as well as the items. They are designed for people who cannot be validly assessed in a language-loaded format — including those with hearing impairment, speech and language disorders, or no shared language with the administrator.

One caveat applies to all of them. “Culture fair” is a design goal, not an achieved state. Familiarity with abstract diagrams, formal schooling and test-taking conventions still varies across populations, and no test has eliminated that entirely.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now! →

Secure & encryptedInstant results10–20 minutes

Group and screening tests

Group-administered cognitive tests are used at scale where individual assessment would be impossible: school-wide ability screening, military and occupational selection, and large research studies. They are efficient and, at the population level, statistically sound.

Their limitation is individual precision. A single low score on a group test can reflect a bad morning, an unfamiliar format, misread instructions or a missed line on an answer sheet, and nothing in the procedure would catch it. Good practice treats a group result as a flag for further assessment, never as a diagnosis. Our article on what makes an IQ test reliable covers why that distinction matters.

Where online IQ tests fit

Online tests occupy a real and legitimate niche, provided everyone is honest about what it is. An unsupervised test cannot verify who is taking it, cannot control the environment, cannot stop you looking something up and cannot prevent retakes. For that reason no online test qualifies anyone for Mensa, and none is diagnostic.

What a well-built online test can do is give you a properly normed estimate. The features that separate a serious one from a novelty quiz are concrete and checkable:

  • It states the scale its score is on, so the number can be compared with anything else.
  • It reports a percentile, not just a score, so the rank is explicit.
  • It gives a confidence range, acknowledging measurement error instead of implying false precision.
  • It is timed and consistent, so your result and someone else’s were produced under the same conditions.
  • It describes its norm group rather than leaving you to guess who you are being compared against.

Our own IQ test reports the scale, percentile and confidence range together, and the IQ tools collection lets you convert a score between scales or see where it sits on the distribution. For a fuller treatment of what separates a good online test from a bad one, see your guide to good online IQ tests.

Choosing the right kind of test

The right test depends entirely on the question you are asking.

  • For a clinical or educational decision — a diagnosis, a placement, an accommodation — only an individually administered test by a qualified professional will do.
  • For someone tested outside their first language, or from a background unlike the norm group, a nonverbal test such as Raven’s gives a fairer reading.
  • For Mensa or another high-IQ society, only an approved supervised test on their published list counts.
  • For personal curiosity, a properly normed online test that reports scale, percentile and confidence range is a reasonable and inexpensive starting point.

One last boundary is worth drawing. None of these tests measures emotional skill, and the emotional intelligence questionnaires sometimes sold alongside them are a different construct with a different evidence base — we compare the two in IQ vs EQ.

What none of them will give you is a single unarguable number that captures your intelligence. Every test samples a slice of cognitive ability through one particular method, and the score is an estimate of where that slice sits relative to other people. Knowing which slice and which method is what turns the number into information.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Tagged accurate iq test, certified iq test, intelligence test, iq assessment, IQ Test, iq test questions, iq testing, iq tests, kids iq test, official iq test, professional iq test, ravens progressive matrices, reliable iq test, standard deviation iq, standard iq test, stanford binet, valid iq test, wechsler adult intelligence scale, what is on an iq test