OpenAI and Anthropic published figures for the same rival models within four days of each other. They do not match, and the reasons are the ones psychometrics named a century ago.
Two frontier model launches landed within four days of each other in early September 2026. Anthropic published Claude Fable 5.1 on 1 September; OpenAI published GPT-6 Astra on 3 September. Because each company benchmarks its rivals as well as itself, the two announcements contain independently produced scores for several of the same models — a rare natural experiment in how reproducible these figures actually are.
They are less reproducible than the decimal places suggest.
Comparing the published figures for models both companies tested:
None of these is an enormous gap, and one of them runs against the direction a cynical reading would predict: OpenAI reports its own model slightly lower than Anthropic does. This is not a story about either company inflating a number. It is a story about what a benchmark score is.
Anthropic states, in its own announcement, that the standard error on Terminal-Bench Science 0.1 is 3.5 to 4.5 points per model. That single disclosure reframes the entire table it appears in. A 3.3-point disagreement between two labs on that benchmark is comfortably inside one standard error — which is to say, not a disagreement at all in any meaningful sense.
It also means that both companies are reporting scores to one decimal place on an instrument whose real precision is roughly plus or minus four points. Independent work points the same way: researchers on the Olmo 3 project found that most post-training evaluations carry standard deviations between 0.25 and 1.5 points even when the evaluation setup is held completely constant, and that changing prompts or sampling parameters moves scores further than that.
A figure quoted to one decimal place, on an instrument accurate to about four points, is not precision. It is decoration.
Readers of this section will recognise the shape of the problem, because it is the one that governs every score a person receives on a cognitive test. A raw result carries measurement error, and a responsible report expresses that as a confidence interval rather than a point — the argument in why an IQ score is a range, not a number. The convention is not a hedge. It is the honest form of the number.
The largest effect in either announcement is not the inter-lab disagreement at all. Anthropic reports Claude Fable 5.1 on OSWorld 2.0 twice, under two different scoring rules: 77.9 per cent under partial scoring, and 41.7 per cent under strict scoring. Same model, same benchmark, same day, 36 points apart.
Neither number is wrong. Partial scoring gives credit for a task substantially completed; strict scoring requires the whole thing. Which one is right depends on what you want to know, and the answer differs by use case. But a headline citing 77.9 per cent and a headline citing 41.7 per cent describe the same performance, and nothing in either figure warns you which convention produced it.
Anthropic adds one more caveat to those numbers that has no clean analogue in human testing and deserves flagging: Fable 5.1 was evaluated with its production safeguards active, and on tasks where those safeguards intervened, the model scored zero. A refusal and a failure are recorded identically. The nearest human parallel is an omitted item scoring the same as a wrong answer — a known artefact that test manuals handle explicitly, precisely because the two mean different things.
OpenAI's footnotes disclose three further methodological choices, each reasonable, each capable of moving a number:
The last of those is the one worth sitting with, and it should be stated neutrally, because it was disclosed openly rather than buried. On an evaluation where a model's answers must be judged, one company's model was used to grade its competitor's output. There is no accusation to make here — automated grading is standard practice and a human panel has its own problems — but it is exactly the question a psychometrician asks about any assessment: who constructed the key, and does the construction favour a particular kind of answer?
Everything above reduces to a checklist that has governed educational and psychological measurement for the better part of a century. A number is a measurement when it has all three, and a figure when it has fewer.
The same reasoning is why this section has argued that national IQ league tables carry far less information than their charts imply. Numbers gathered under conditions that vary between entries can be tabulated, ranked and coloured on a map, and the ranking will still not mean what the reader assumes.
The practical version is one question, and it works on model announcements and online IQ results alike: under what conditions was this produced, and what is the error around it? A score that cannot answer both is not yet a measurement.
Find your IQ score now! →None of this is an argument that AI benchmarks are worthless. They are the best available instrument for a genuinely hard measurement problem, and both companies published enough detail for an outsider to reconstruct the discrepancies — which is more transparency than most industries offer. The argument is narrower: the disclosure lives in the footnotes and the number travels alone. Anyone comparing figures across two announcements is comparing results produced under different conditions, and should treat gaps smaller than a few points as noise until told otherwise.
Because the evaluation setup is not standardised across labs. Companies use different test harnesses, different grading rules, different prompts and sampling parameters, and sometimes different graders. Anthropic reports Claude Fable 5 at 24.7 per cent on Terminal-Bench Science 0.1 where OpenAI reports 21.4 per cent; the gap is inside the 3.5 to 4.5 point standard error Anthropic states for that benchmark.
Only with large error bars. The evaluation process each company uses internally is not controlled across models or fully documented, so figures from separate press releases are not directly comparable. Where both companies publish identical figures for the same model — as they do for Humanity's Last Exam — the comparison is sounder.
That the figure shown is the best result across the model's reasoning-effort settings rather than a result at a fixed compute budget. It is disclosed at the foot of OpenAI's September 2026 comparison table. The human equivalent would be reporting a candidate's best score across several sittings without stating the conditions of any of them.
It varies by benchmark and is rarely printed. Anthropic states 3.5 to 4.5 points per model on Terminal-Bench Science 0.1. Independent work on the Olmo 3 project found standard deviations of 0.25 to 1.5 points even with the evaluation setup held constant, with larger movements from changes to prompts or sampling. Most published figures carry one decimal place and no interval.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
Every properly built test publishes a standard error of measurement alongside the score. On a mean-100, standard-deviation-15 scale it is large enough that a reported 104 and a reported 111 are not reliably different results.
The tables that circulate as national IQ figures come from one compilation built out of whatever studies existed, in whatever years, on whatever samples. Several of the numbers were estimated rather than measured.

Nearly every "IQ by profession" chart online traces to one 1945 study of wartime enlisted men, scored on a scale that is not even the one modern IQ tests use. Here is the real table — and what still holds up.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →