IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment
Data & RankingsWell established

The Same Model, Two Labs, Different Scores: Why AI Benchmarks Disagree

OpenAI and Anthropic published figures for the same rival models within four days of each other. They do not match, and the reasons are the ones psychometrics named a century ago.

Illustration generated for IQ Metrics. No photograph is used. IQ Metrics

What you need to know

  • Anthropic published Claude Fable 5.1 figures on 1 September and OpenAI published GPT-6 Astra figures on 3 September. Both tables report scores for the same Claude models, and on several benchmarks the two companies disagree by more than a point.
  • On Terminal-Bench Science 0.1, OpenAI reports Claude Fable 5 at 21.4 per cent and Anthropic reports 24.7 per cent. Anthropic separately states the standard error on that benchmark is 3.5 to 4.5 points per model — wider than most of the gaps being argued over, and wider than the decimal places imply.
  • The largest single effect is not between labs but between grading rules. Anthropic reports Fable 5.1 on OSWorld 2.0 at 77.9 per cent under partial scoring and 41.7 per cent under strict scoring: the same model, the same benchmark, 36 points apart.
  • Both companies disclose their methods in footnotes. The problem is not concealment. It is that a headline number carries none of the disclosure with it.

Two frontier model launches landed within four days of each other in early September 2026. Anthropic published Claude Fable 5.1 on 1 September; OpenAI published GPT-6 Astra on 3 September. Because each company benchmarks its rivals as well as itself, the two announcements contain independently produced scores for several of the same models — a rare natural experiment in how reproducible these figures actually are.

They are less reproducible than the decimal places suggest.

Where the two tables disagree

Comparing the published figures for models both companies tested:

  • Terminal-Bench Science 0.1, Claude Fable 5 — OpenAI reports 21.4 per cent, Anthropic reports 24.7 per cent.
  • Terminal-Bench Science 0.1, Claude Opus 5 — OpenAI 30.0 per cent, Anthropic 29.0 per cent.
  • Terminal-Bench 4.0, Claude Fable 5 — OpenAI 44.5 per cent, Anthropic 42.0 per cent.
  • AutomationBench, GPT-5.6 Sol — OpenAI reports its own model at 18.1 per cent, Anthropic reports it at 19.6 per cent.
  • OSWorld 2.0, Claude Opus 5 — OpenAI 70.2 per cent, Anthropic 75.4 per cent.

None of these is an enormous gap, and one of them runs against the direction a cynical reading would predict: OpenAI reports its own model slightly lower than Anthropic does. This is not a story about either company inflating a number. It is a story about what a benchmark score is.

The error bar nobody prints next to the figure

Anthropic states, in its own announcement, that the standard error on Terminal-Bench Science 0.1 is 3.5 to 4.5 points per model. That single disclosure reframes the entire table it appears in. A 3.3-point disagreement between two labs on that benchmark is comfortably inside one standard error — which is to say, not a disagreement at all in any meaningful sense.

It also means that both companies are reporting scores to one decimal place on an instrument whose real precision is roughly plus or minus four points. Independent work points the same way: researchers on the Olmo 3 project found that most post-training evaluations carry standard deviations between 0.25 and 1.5 points even when the evaluation setup is held completely constant, and that changing prompts or sampling parameters moves scores further than that.

A figure quoted to one decimal place, on an instrument accurate to about four points, is not precision. It is decoration.

Readers of this section will recognise the shape of the problem, because it is the one that governs every score a person receives on a cognitive test. A raw result carries measurement error, and a responsible report expresses that as a confidence interval rather than a point — the argument in why an IQ score is a range, not a number. The convention is not a hedge. It is the honest form of the number.

The grading rule matters more than the lab

The largest effect in either announcement is not the inter-lab disagreement at all. Anthropic reports Claude Fable 5.1 on OSWorld 2.0 twice, under two different scoring rules: 77.9 per cent under partial scoring, and 41.7 per cent under strict scoring. Same model, same benchmark, same day, 36 points apart.

Neither number is wrong. Partial scoring gives credit for a task substantially completed; strict scoring requires the whole thing. Which one is right depends on what you want to know, and the answer differs by use case. But a headline citing 77.9 per cent and a headline citing 41.7 per cent describe the same performance, and nothing in either figure warns you which convention produced it.

Anthropic adds one more caveat to those numbers that has no clean analogue in human testing and deserves flagging: Fable 5.1 was evaluated with its production safeguards active, and on tasks where those safeguards intervened, the model scored zero. A refusal and a failure are recorded identically. The nearest human parallel is an omitted item scoring the same as a wrong answer — a known artefact that test manuals handle explicitly, precisely because the two mean different things.

Who administers the test, and who writes the answer key

OpenAI's footnotes disclose three further methodological choices, each reasonable, each capable of moving a number:

  • On OSWorld 2.0, Claude's scores use "the official settings, and not the modified tasks and modified grading" from Anthropic's own system card — an explicit statement that the two companies ran the benchmark differently.
  • On BenchCAD, Claude's scores "reflect 3 modifications to the eval" documented in Anthropic's system card.
  • On HealthBench Professional, OpenAI independently evaluated all Claude models using GPT-5.4 as the grader.

The last of those is the one worth sitting with, and it should be stated neutrally, because it was disclosed openly rather than buried. On an evaluation where a model's answers must be judged, one company's model was used to grade its competitor's output. There is no accusation to make here — automated grading is standard practice and a human panel has its own problems — but it is exactly the question a psychometrician asks about any assessment: who constructed the key, and does the construction favour a particular kind of answer?

A fourth footnote records that Claude Fable 5 and 5.1 were excluded from three life-sciences evaluations "because they refuse the majority of questions". That is a real and legitimate reason to omit a model. It also means an absent row can reflect a safety configuration rather than a capability limit, and the table cannot show you which.

The three conditions a score has to meet

Everything above reduces to a checklist that has governed educational and psychological measurement for the better part of a century. A number is a measurement when it has all three, and a figure when it has fewer.

  • A defined reference population. A score is a position among some group; without the group there is no position. Machine evaluation has no population and reports raw percentages instead, which is why a benchmark result cannot be converted into an IQ.
  • A norming sample, meaning the instrument administered under fixed conditions to a representative slice of that population.
  • Standardised administration, so that two results mean the same thing. This is the one the September tables most visibly lack: different harnesses, different grading rules, different graders, and a disclosed policy of reporting "the maximum at any effort" rather than a result at a fixed budget.

The same reasoning is why this section has argued that national IQ league tables carry far less information than their charts imply. Numbers gathered under conditions that vary between entries can be tabulated, ranked and coloured on a map, and the ranking will still not mean what the reader assumes.

Your own number

Where would your own score land?

The practical version is one question, and it works on model announcements and online IQ results alike: under what conditions was this produced, and what is the error around it? A score that cannot answer both is not yet a measurement.

Find your IQ score now!
Secure & encryptedInstant results10–20 minutes

None of this is an argument that AI benchmarks are worthless. They are the best available instrument for a genuinely hard measurement problem, and both companies published enough detail for an outsider to reconstruct the discrepancies — which is more transparency than most industries offer. The argument is narrower: the disclosure lives in the footnotes and the number travels alone. Anyone comparing figures across two announcements is comparing results produced under different conditions, and should treat gaps smaller than a few points as noise until told otherwise.

Common questions

Why do AI benchmark scores differ between companies?

Because the evaluation setup is not standardised across labs. Companies use different test harnesses, different grading rules, different prompts and sampling parameters, and sometimes different graders. Anthropic reports Claude Fable 5 at 24.7 per cent on Terminal-Bench Science 0.1 where OpenAI reports 21.4 per cent; the gap is inside the 3.5 to 4.5 point standard error Anthropic states for that benchmark.

Can I compare numbers from two different model announcements?

Only with large error bars. The evaluation process each company uses internally is not controlled across models or fully documented, so figures from separate press releases are not directly comparable. Where both companies publish identical figures for the same model — as they do for Humanity's Last Exam — the comparison is sounder.

What does "maximum at any effort" mean on a benchmark table?

That the figure shown is the best result across the model's reasoning-effort settings rather than a result at a fixed compute budget. It is disclosed at the foot of OpenAI's September 2026 comparison table. The human equivalent would be reporting a candidate's best score across several sittings without stating the conditions of any of them.

How large is the error on an AI benchmark score?

It varies by benchmark and is rarely printed. Anthropic states 3.5 to 4.5 points per model on Terminal-Bench Science 0.1. Independent work on the Olmo 3 project found standard deviations of 0.25 to 1.5 points even with the evaluation setup held constant, with larger movements from changes to prompts or sampling. Most published figures carry one decimal place and no interval.

Sources for this story

  1. GPT-6 Astra: A new generation of intelligence, comparison tables and footnotes 3, 5, 11 and 12 — OpenAI
  2. Introducing Claude Fable 5.1 and Claude Mythos 5.1, benchmark tables, the stated standard error and the safeguards caveat — Anthropic
  3. Olmo 3 project findings on evaluation variance under a fixed setup — Allen Institute for AI
  4. BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices — arXiv
  5. Standards for Educational and Psychological Testing, on norming, standardised administration and score reporting — American Educational Research Association, American Psychological Association and National Council on Measurement in Education

Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.

Share this story

Know someone who keeps seeing this number quoted without the scale it was measured on? Send it to them — it takes one tap.

Filed under#ai benchmarks#study quality#confidence intervals#test conditions

Read the research.
Then find your own number.

Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.

Start IQ Test
Secure & encryptedInstant results10–20 minutes