Professional IQ Tests: What Psychologists Actually Use
When a psychologist measures intelligence, they reach for one of a small number of standardised instruments. This guide covers what the Wechsler scales, Stanford-Binet, Woodcock-Johnson and Raven’s Progressive Matrices each measure, who each one is for, and where an online test genuinely fits.

“IQ test” is a category, not a product. When a psychologist actually measures intelligence, they reach for one of a small number of standardised instruments, and those instruments differ in how they are delivered, what abilities they sample, how long they take, what age group they were normed on, and how heavily they lean on language. A score means very little until you know which of them produced it.
This guide walks through the tests professionals actually use, what each is genuinely good at, and how to work out which kind of result you are looking at. If you want to compare the shorter formats available on this site instead, the IQ tests hub covers those.
The first division: individual or group
The most important split is not verbal versus nonverbal — it is whether a trained administrator sits with you.
Individually administered tests are delivered one-to-one over roughly one to two hours. The administrator controls timing, watches how you approach each item, and can tell the difference between not knowing an answer and misunderstanding the question. These are the tests used clinically, educationally and diagnostically, and they are the reference standard for accuracy.
Group tests are delivered to many people at once, usually on paper or screen, with fixed instructions and no individual observation. They are far cheaper and faster, which is why they dominate schools, military selection and large-scale screening. They are also less precise for any single individual, because none of the observational detail survives.
The major individually administered tests

The Wechsler scales: WAIS and WISC
The Wechsler tests are the most widely used individual intelligence tests in the world. The WAIS covers adults and older adolescents; the WISC covers school-age children; the WPPSI covers preschoolers. They are periodically revised, and each revision is renormed on a fresh standardisation sample.
Rather than producing one number, a Wechsler test reports a full-scale score built from index scores covering verbal comprehension, visual-spatial and fluid reasoning, working memory and processing speed. That structure is the real value: two people can reach the same composite by very different routes, and the profile of indexes is often more informative than the headline figure. Wechsler tests use a standard deviation of 15.
The Stanford-Binet Intelligence Scales
The Stanford-Binet is the direct descendant of the original Binet-Simon scale and the oldest continuously developed intelligence test still in use. Its current edition covers an unusually wide age span in a single instrument and assesses five factors in both verbal and nonverbal form, which makes it useful at the extremes of the range where other tests run out of items. It uses a standard deviation of 16, which is why Stanford-Binet scores sit slightly higher than Wechsler scores for identical performance — a difference our IQ score converter handles directly. The test’s long history is covered in our piece on the history of IQ testing.
The Woodcock-Johnson Tests of Cognitive Abilities
The Woodcock-Johnson is built explicitly on the Cattell-Horn-Carroll model of intelligence, which organises cognitive ability into a hierarchy of broad and narrow factors. Its distinguishing feature is that it pairs with a co-normed achievement battery, so cognitive ability and academic attainment can be compared on the same standardisation sample. That makes it a common choice in educational assessment, particularly where a learning difficulty is in question.
The Kaufman batteries
The KABC-II was designed with a specific concern in mind: reducing the influence of acquired knowledge and language on the resulting score. It offers alternative theoretical scoring models and a version that minimises verbal demands, which makes it useful for children from varied linguistic and cultural backgrounds.
Nonverbal and culture-reduced tests
Every verbal test carries an unavoidable problem: it measures language proficiency alongside reasoning. For someone tested in a second language, or from a background unlike the norm group, that confound can be large. Nonverbal tests exist to reduce it.
Raven’s Progressive Matrices
Raven’s is the best known nonverbal reasoning test and one of the purest measures of fluid intelligence available. Every item is a visual pattern with a piece missing; you choose the piece that completes the rule. There are no words, no cultural references and no general knowledge anywhere in it. It comes in progressive difficulty levels, from Coloured for young children through Standard to Advanced for high-ability adults.
The Cattell Culture Fair Intelligence Test
The CFIT was built for the same purpose using series completion, classification, matrix and conditional reasoning items, all nonverbal. It reports on a scale with a standard deviation of 24, which is why Cattell scores look dramatically higher than Wechsler ones — the ruler is wider, not the person. This is the scale British Mensa uses, as covered in our article on Mensa IQ score requirements.
TONI, UNIT and Leiter
These go further still, minimising or removing language from the instructions as well as the items. They are designed for people who cannot be validly assessed in a language-loaded format — including those with hearing impairment, speech and language disorders, or no shared language with the administrator.
One caveat applies to all of them. “Culture fair” is a design goal, not an achieved state. Familiarity with abstract diagrams, formal schooling and test-taking conventions still varies across populations, and no test has eliminated that entirely.
Where would your own score land?
Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.
Group and screening tests
Group-administered cognitive tests are used at scale where individual assessment would be impossible: school-wide ability screening, military and occupational selection, and large research studies. They are efficient and, at the population level, statistically sound.
Their limitation is individual precision. A single low score on a group test can reflect a bad morning, an unfamiliar format, misread instructions or a missed line on an answer sheet, and nothing in the procedure would catch it. Good practice treats a group result as a flag for further assessment, never as a diagnosis. Our article on what makes an IQ test reliable covers why that distinction matters.
Where online IQ tests fit
Online tests occupy a real and legitimate niche, provided everyone is honest about what it is. An unsupervised test cannot verify who is taking it, cannot control the environment, cannot stop you looking something up and cannot prevent retakes. For that reason no online test qualifies anyone for Mensa, and none is diagnostic.
What a well-built online test can do is give you a properly normed estimate. The features that separate a serious one from a novelty quiz are concrete and checkable:
- It states the scale its score is on, so the number can be compared with anything else.
- It reports a percentile, not just a score, so the rank is explicit.
- It gives a confidence range, acknowledging measurement error instead of implying false precision.
- It is timed and consistent, so your result and someone else’s were produced under the same conditions.
- It describes its norm group rather than leaving you to guess who you are being compared against.
Our own IQ test reports the scale, percentile and confidence range together, and the IQ tools collection lets you convert a score between scales or see where it sits on the distribution. For a fuller treatment of what separates a good online test from a bad one, see your guide to good online IQ tests.
Choosing the right kind of test
The right test depends entirely on the question you are asking.
- For a clinical or educational decision — a diagnosis, a placement, an accommodation — only an individually administered test by a qualified professional will do.
- For someone tested outside their first language, or from a background unlike the norm group, a nonverbal test such as Raven’s gives a fairer reading.
- For Mensa or another high-IQ society, only an approved supervised test on their published list counts.
- For personal curiosity, a properly normed online test that reports scale, percentile and confidence range is a reasonable and inexpensive starting point.
One last boundary is worth drawing. None of these tests measures emotional skill, and the emotional intelligence questionnaires sometimes sold alongside them are a different construct with a different evidence base — we compare the two in IQ vs EQ.
What none of them will give you is a single unarguable number that captures your intelligence. Every test samples a slice of cognitive ability through one particular method, and the score is an estimate of where that slice sits relative to other people. Knowing which slice and which method is what turns the number into information.
Keep reading
Taking a Test
What Are The Type of Questions Asked On IQ Test?
Remember that IQ tests only evaluate cognitive power momentarily. Maintaining a growth mindset and viewing exams as learning opportunities is key.
Scores & Scales
Mensa IQ Score Requirements
Mensa admits the top 2% of scorers, but "the top 2%" is a rank, not a number, and the number it maps to changes with the test you sat. Here is what qualifies on the Wechsler, Stanford-Binet and Cattell scales, the two routes to membership, and what the supervised admission test is actually like.
Taking a Test
What Makes IQ Test Reliable ?
Many academics, teachers, and psychologists are interested in and have debated this subject. In “The Science Behind Intelligence Testing: Unveiling the Reliability of IQ Tests,” we examine the complex field of IQ testing and the scientific evidence supporting its validity.