What Is a Good Watson-Glaser Score? What the Test Measures
Pearson publishes no pass mark and says a raw score out of 40 cannot be called good without a norm group. It does publish 12, 12 and 16 items in three sections, and a .52 correlation with WAIS-IV Full Scale IQ.

What you need to know
- Pearson's efficacy report says that without a norm we cannot know whether a score is good. Its worked example: a raw 35 out of 40 beats about 86 per cent of managers. It publishes no pass mark and advises employers to review pass rates by group before choosing one.
- The current Watson-Glaser III gives each test-taker 40 items from a bank: 12 on recognising assumptions, 12 on evaluating arguments and 16 on drawing conclusions. Pearson's summary gives a 30-minute limit, and Clifford Chance says 40 questions in 30 minutes.
- Its link with IQ is measurable: in the Watson-Glaser II manual, totals correlated .52 with WAIS-IV Full Scale IQ in 56 adults (mean 110.8, SD 14.6). Pearson's efficacy report gives .53 with Raven's Advanced Progressive Matrices.
- Prep-site targets (70th percentile, 33 to 35 out of 40, 80 per cent) cite no source. On the only score distribution we could open, a 2009 standardisation sample of 636, 33 to 35 out of 40 is the 82nd to 89th percentile, not the 70th.
There is no single good Watson-Glaser score. Pearson, which publishes the test, says a raw score out of 40 cannot be judged good or bad without a norm group, and it publishes no pass mark. What it does publish is a good deal about how the test is built: 40 items in three sections, a percentile against one of 18 norm groups, an internal consistency of .83, and a correlation of .52 with the WAIS-IV Full Scale IQ. This guide sets out those figures, what two law firms say about the test, and where the good-score numbers on prep sites come from.
What is the Watson-Glaser test?
The Watson-Glaser Critical Thinking Appraisal is a test of critical thinking used in graduate and professional hiring, especially by law firms. Pearson says it was first published in 1964 by Goodwin Watson and Edward M. Glaser, after development dating back to 1926. The timeline in Pearson's 2020 efficacy report runs from two parallel 100-item forms in 1964, to two 80-item forms in 1980, a UK adaptation in 1991 and a 40-item short form in 1994. The Watson-Glaser II arrived in 2010 with two 40-item forms. The current Watson-Glaser III, which replaced fixed forms with an item bank, was released in the UK in 2012 and in the US in 2018. It is completed online and suits supervised or unsupervised use.
How many questions is it, and how long does it take?
Each test-taker receives 40 items drawn from a large bank. Pearson's one-page efficacy summary gives a 30-minute time limit and says an untimed US-English version exists for development or reasonable accommodations. Clifford Chance's London careers page agrees: the test includes 40 questions and lasts 30 minutes, which is 45 seconds an item, and candidates have four days including weekends to complete it. Linklaters gives candidates five days from the date of application. Pearson's full efficacy report calls the recommended limit intentionally generous without stating minutes, and a 2023 French catalogue sheet gives 20 minutes, so check the time on your own invitation.
What are the three sections?
Pearson says the test's structure follows the RED model, which breaks critical thinking into recognising assumptions, evaluating arguments and drawing conclusions. Each test-taker is presented with 40 items split across those three.
- Recognise assumptions (12 items): spotting what a statement takes for granted without saying so
- Evaluate arguments (12 items): judging how strong an argument is and whether it bears on the question
- Draw conclusions (16 items): deciding what follows from the information given
Older material uses five subtests: inference, recognition of assumptions, deduction, interpretation and evaluation of arguments. Pearson says factor analysis showed inference, deduction and interpretation grouping together, so since the Watson-Glaser II of 2010 they have been reported as one, draw conclusions. Clifford Chance's page still lists the five older skills, and Pearson's UK practice booklet keeps the five-part layout. Our guide to syllogism questions covers the kind of reasoning the conclusion items ask for.
How is a Watson-Glaser score reported?
The raw score is the number correct out of 40, converted to a percentile against a norm group. Pearson's worked example: a raw 35 out of 40 is better than about 86 per cent of managers when set against its Manager norm. It then makes the point that matters here: without applying a norm, we cannot know whether a score is good or not. The current edition comes with 18 norms built by Pearson, by occupation, such as accountant, engineer and HR professional, and by organisational level from entry level to executive. Pearson says an employer building its own norm should ideally have at least 200 test-takers. It gives no sizes for the current edition's norms.
The 2009 Watson-Glaser II manual does list sizes for its normative samples, which shows the scale involved: 3,243 managers, 1,468 directors, 1,389 executives, 368 accountants and 306 hourly or entry-level workers, an average of 967 across 14 groups. Those belong to the older edition. The manual also gives the one score distribution we could open: in its standardisation sample of 636 people, Form D totals averaged 27.1 out of 40 with a standard deviation of 6.5. On a normal curve that puts a raw score at the percentiles below. This is our arithmetic on a 2009 sample of a different edition, not a Pearson conversion table.
- 20 of 40: about the 14th percentile
- 24 of 40: about the 32nd
- 27 of 40: about the 49th
- 30 of 40: about the 67th
- 32 of 40: about the 77th
- 33 of 40: about the 82nd
- 35 of 40: about the 89th
- 38 of 40: about the 95th
What is a good Watson-Glaser score?
Pearson publishes no pass mark. Its efficacy report says that even where analyses find no evidence of statistical bias in a test, differences in actual pass rates between groups can be reduced by careful selection of pass marks, and it advises customers to review pass rates for different groups and choose a pass mark with minimal differences. The bar is the employer's. Clifford Chance's London page and Linklaters' early-careers page both describe the test and its deadline, and neither states a pass mark or says how the score is combined with other stages.
The good-score numbers on prep sites have no source we could find. The ones we read say the 70th percentile or above, roughly 33 to 35 out of 40, is needed for competitive firms; that the average is about 55 per cent and 80 per cent or more is the aim; or that 80 to 90 per cent is typically required by the Magic Circle firms, equating 32 of 40 with 80 per cent. None cites a Pearson or firm document. A percentage of items correct is not a percentile, and the two disagree here: on the 2009 sample above, 33 to 35 out of 40 is the 82nd to 89th percentile, not the 70th. Different edition, different sample, so treat all of these figures as loose.
Pearson's own answer to what a good score is: without a norm group nobody can say, and the pass mark belongs to the employer.
Is the Watson-Glaser an IQ test?
No, but it overlaps with IQ more than most hiring tests can show. The Watson-Glaser II manual reports that Form D total scores correlated .52 with the WAIS-IV Full Scale IQ in 56 adults with a bachelor's degree or higher. That group averaged 110.8 on the WAIS-IV with a standard deviation of 14.6, on the scale where the general norm is a mean of 100 and a standard deviation of 15. Against the WAIS-IV indexes the correlations were .44 for working memory, .46 for perceptual reasoning and .42 for verbal comprehension, and .14, not significant, for processing speed. Pearson's efficacy report adds .53 with Raven's Advanced Progressive Matrices and .47 with a numerical data interpretation test in 91 US working adults.
A correlation of .52 means the two measures share about 27 per cent of their variation, its square. With only 56 people the margin is wide: our calculation puts the 95 per cent interval at roughly .30 to .69. So the test sits in the same family as IQ measures without being one, and Pearson reports no conversion. Our guides to WAIS index scores and Raven's matrices explain the two comparison tests. The Watson-Glaser also served as one of the two intelligence measures in a 900-person study of whether personality predicts IQ, which our guide to MBTI type and IQ covers.
How reliable are the scores?
Pearson's efficacy report gives a Cronbach's alpha, a measure of whether the items hang together, of .83 for the current edition in 147 US working adults, and .75 to .86 for earlier versions. Test-retest correlations across versions run from .73 to .89, and two UK studies of alternate forms, with 355 and 318 people, gave .82 and .88. Section scores are weaker. In the Watson-Glaser II manual the total-score alpha was .83 in 1,011 people, but the sections gave .80 for recognising assumptions, .70 for drawing conclusions and .57 for evaluating arguments, which the manual itself calls low. That is a reason to treat the total as the result and the three section percentiles as rough guides.
How should you prepare, and what does an employer do with the result?
Pearson's UK edition has a practice booklet, and Linklaters strongly advises a practice test. Practice helps with the format: our guide to the practice effect puts a second sitting at around a quarter of a standard deviation higher on average. On the employer side, cut scores are expected to be checked for group differences, as our guides to whether IQ tests are legal for hiring and the four-fifths rule explain. Other graduate employers use different tests with the same logic; our guides to the CCAT, the Wonderlic and SHL scores show how each reports a result.
Reasoning well is not the same as reasoning without error. Our guide to why smart people make bad decisions covers what a critical thinking score does not capture, and our verbal IQ test covers the reading-based reasoning that the assumption and argument items lean on.
Where would your own score land?
Want a result on the standard IQ scale, with its percentile? Take the IQ Metrics test.
Find your IQ score now! →A Watson-Glaser result is a rank within a norm group chosen by the employer, plus a raw count out of 40. The IQ Metrics test reports a score on the standard IQ scale with its percentile, and the percentile calculator shows how any score sits on the curve. Our guide to what a cognitive test score predicts at work sets the wider evidence around tests like this one. Another common employer test, the Predictive Index Cognitive Assessment, is covered in our guide to PI scores.
Common questions
What is a good Watson-Glaser score?
Pearson publishes no pass mark and says a raw score out of 40 cannot be judged good without a norm group. Employers set their own cut-offs, and neither Clifford Chance nor Linklaters states one. Figures such as the 70th percentile or 33 to 35 out of 40 come from prep sites that cite no source.
How many questions are on the Watson-Glaser test and how long does it take?
The current edition has 40 items. Pearson's summary gives a 30-minute limit, and Clifford Chance's London page says 40 questions in 30 minutes. A 2023 French catalogue sheet gives 20 minutes, so check your invitation.
What are the sections of the Watson-Glaser test?
Three, following Pearson's RED model: recognise assumptions (12 items), evaluate arguments (12 items) and draw conclusions (16 items). Older material used five subtests: inference, recognition of assumptions, deduction, interpretation and evaluation of arguments.
Is the Watson-Glaser an IQ test?
No, but it overlaps with IQ. In the Watson-Glaser II manual, totals correlated .52 with WAIS-IV Full Scale IQ in 56 adults, and Pearson's efficacy report gives .53 with Raven's Advanced Progressive Matrices. Pearson publishes no conversion to an IQ score.
Sources for this story
- Assessment Efficacy Report: Watson-Glaser Critical Thinking Appraisal III (April 2020) — Pearson
- Watson-Glaser Critical Thinking Appraisal: Efficacy Report Summary — Pearson
- Watson-Glaser II Critical Thinking Appraisal Technical Manual and User's Guide (2009) — NCS Pearson
- Watson-Glaser Critical Thinking Appraisal, UK Edition: practice test booklet — Pearson Assessment
- How we hire: London, the Watson-Glaser test — Clifford Chance
- Your application: early careers and the Watson-Glaser test — Linklaters
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
Related stories
All news →
What Is a Good CCAT Score? Average, Range and the IQ Question
Criteria publishes a median of 24 out of 50 and suggested minimums that run from 17 to 24 by job. The 30-plus targets and IQ tables that circulate online appear on prep sites, not in the Criteria documents we read.

Syllogism Questions on IQ Tests: Why the Right Answer Feels Wrong
A syllogism gives you two statements and asks what must follow. Decades of research show most people answer with what they already believe instead — and the fix is one simple substitution trick.

What Is a Good Predictive Index Cognitive Assessment Score?
PI scores run from 100 to 450, and its own blog says the average of 250 is about 20 of 50 questions right. PI says it is not a test you pass or fail, no longer shows percentiles, and does not measure IQ.
Read the research.
Then find your own number.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →
