GPT-6 Astra scored 62.7 per cent and 99.9 per cent on the same benchmark in the same week. The difference is not the model. It is how the test was administered.
In the first week of September 2026 the same model was reported as scoring 62.7 per cent and 99.9 per cent on the same benchmark. Both numbers are correct. Both were published by the same organisation, on the same day, about the same system. The gap between them is almost entirely a fact about how the test was administered, and it is the most instructive thing to come out of the launch.
ARC-AGI-3, announced in March 2026, is the third generation of a benchmark series built around a specific argument: that most AI evaluations measure skill at a task, and skill at a task can be bought with enough training data. What the series tries to measure instead is how efficiently a system picks up something it has never been trained on.
The third version is the first that is interactive. A system is dropped into a small turn-based environment with no instructions and has to work out the goal, the controls and the rules by acting inside it. Every environment is calibrated on people before release, and humans solve 100 per cent of them — a design requirement rather than a result, and the detail that makes a low machine score meaningful. We set out that calibration in full in our earlier report on what ARC-AGI-3 measures and what it does not.
A harness is the scaffolding around the model: what it is shown each turn, what it is allowed to carry forward, how its previous thinking is stored. The ARC Prize Foundation now runs two.
Under the Standard harness, Astra's best result was 62.7 per cent, at maximum reasoning effort, for 26,098 dollars of compute. Under the Provider Adapter, its best was 99.9 per cent, at high reasoning effort, for 18,817 dollars.
Published alongside those headline figures is a full breakdown by reasoning effort, and one line in it is worth more than the rest combined. At the lowest setting — reasoning effort switched off entirely — the two harnesses diverge like this:
With the model's deliberate reasoning turned off, the provider harness still returns a result close to its ceiling. Whatever is producing the 99.9 per cent, it is not primarily the reasoning the benchmark was built to measure. It is the machinery around the model that carries state from one move to the next.
With reasoning effort switched off, one harness scores 35.2 per cent and the other 96.7 per cent. That is a fact about the scaffolding, not the intelligence.
The Standard column has a second oddity worth noting without over-reading: it is not monotonic. The "low" setting returns 17.5 per cent, below the 35.2 per cent recorded with no reasoning effort at all. Benchmarks at this scale carry real run-to-run variation, and a single non-monotonic step is a reminder that individual cells in these tables are noisier than their decimal places suggest.
There is also a genuinely counterintuitive result in the cost column: higher reasoning effort is cheaper, not dearer. Maximum effort under the Standard harness cost 26,098 dollars; no effort cost 49,791. The model that thinks harder solves games in fewer moves, and fewer moves means fewer calls.
OpenAI's announcement prints a single ARC-AGI-3 row. Astra appears at 99.9 per cent, next to 30.2 per cent for Claude Opus 5 and 7.8 per cent for GPT-5.6 Sol. A footnote records that Astra was run with OpenAI's own responses API harness, and states that the settings changed "do not specifically target ARC-AGI-3". The two comparison models were not run that way.
Discounting the headline number is not the same as discounting the result, and the underlying achievement is substantial. Before launching ARC-AGI-3 the foundation tested roughly 500 members of the general public, not selected for puzzle-solving ability, to establish a baseline for action efficiency — the number of moves a solver needs to work an environment out. Under the provider harness, Astra used fewer actions than the median tested human on 96.0 per cent of the levels it completed, and 51.7 per cent fewer actions per level on average.
The foundation calls that a material milestone, and it is. It is also a specific claim about efficiency in moves, not about ability in general and not about cost. On cost the comparison runs the other way by several orders of magnitude: human participants were paid 115 dollars for a ninety-minute session plus 5 dollars per completed game, working out at roughly 12.78 dollars per game attempted, against the 17,000 to 26,000 dollars a full model run consumed.
The evaluators also noted a caveat about testing conditions that anyone who works with standardised assessment will recognise. In a separate, more permissive harness the model was given a sandbox and wrote its own tools — a maze solver, a combat model, a patrol tracker. The foundation observed that its human participants "did not have a code interpreter, scratch pad, etc.", so those results describe the model and its tools together rather than the model alone.
Two scores thirty-seven points apart, differing only in how the test was administered, is a problem cognitive assessment solved a long time ago and solved strictly. Standardised administration — the same instructions, the same time limit, the same permitted materials — is not an administrative courtesy wrapped around the instrument. It is part of the instrument. A test given under different conditions is a different test, and its norms no longer apply.
That is why a proctored assessment specifies what a candidate may bring into the room, why an untimed administration of a timed test produces a score that cannot be read against the published tables, and why the margin of error around any single figure is reported rather than hidden — the argument we make in why a score is a range and not a number. Machine evaluation is arriving at the same conclusions from first principles, in public, at speed.
The transferable habit is small and works everywhere: before believing a score, ask what it was produced under. Our own test reports a result as a position in a reference population, under stated conditions, with a range around it — the three things a bare percentage cannot supply.
Find your IQ score now! →One last thing deserves quoting directly, because it came from the organisation with the most to gain from the opposite claim. Asked whether the result represents artificial general intelligence, the ARC Prize Foundation wrote that while it believes the system represents meaningful progress towards generalisation, it is "not claiming that it is AGI" — and noted that it had said at launch that saturating the benchmark would not constitute proof of achieving it. When the benchmark's own authors decline the headline, that is worth more than the headline.
The third generation of the Abstraction and Reasoning Corpus benchmark, announced in March 2026 and the first that is interactive. A system is placed in a small turn-based environment with no instructions and must work out the goal, controls and rules by acting. Every environment is calibrated on human participants before release, and humans solve 100 per cent of them.
Different test harnesses. The 62.7 per cent came from the ARC Prize Foundation's provider-neutral harness, which leaves the model to decide what to record in its own visible notes. The 99.9 per cent came from a harness supplied by OpenAI that preserves the model's internal reasoning state between moves and compacts long conversations so earlier work can be reused. The model and the benchmark were the same in both cases.
For comparing models against each other, the Standard harness figure, because every system is run the same way. For describing what a system can do when used as its developer intends, the Provider Adapter figure. Quoting one model's adapter score against another model's standard score is not a comparison. The ARC Prize Foundation now labels both conditions separately on its leaderboard.
The organisation that built the benchmark says not. It has stated that it is not claiming the result is AGI, and that it said at the benchmark's launch that saturating it would not represent proof of achieving AGI. It defines AGI as the ability to acquire any skill a human can, as efficiently as a human can.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
Frontier systems now clear reasoning benchmarks that defeated them a year ago. Put them in front of puzzles screened so that ordinary people solve every one, and the best models finish under one per cent.
IQ scales are defined by a human reference sample, and a language model is not in it. Placing a machine on that scale is not a hard measurement problem — it is a category error.
Machines have been writing reasoning items since long before the current wave of language models, and the research is clear about where the difficulty lies. Generating a plausible question is the easy half. Knowing how hard it is remains the expensive half.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →