IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

Two Scores, One Model: How the Harness Decides an ARC-AGI-3 Result

GPT-6 Astra scored 62.7 per cent and 99.9 per cent on the same benchmark in the same week. The difference is not the model. It is how the test was administered.

Illustration generated for IQ Metrics. No photograph is used. IQ Metrics

What you need to know

  • On ARC-AGI-3, the ARC Prize Foundation measured GPT-6 Astra at 62.7 per cent under its own provider-neutral harness and 99.9 per cent under a harness supplied by OpenAI, which preserves the model's reasoning state between moves.
  • The clearest evidence that the harness is doing the work: with reasoning effort switched off entirely, the provider harness still scores 96.7 per cent while the neutral harness scores 35.2 per cent.
  • OpenAI's announcement prints the 99.9 per cent figure next to 30.2 per cent for Claude Opus 5 and 7.8 per cent for GPT-5.6 Sol. A footnote discloses that Astra used the provider harness; the comparison models did not.
  • The ARC Prize Foundation, whose benchmark this is, states plainly that it is not claiming the result is AGI, and that saturating the benchmark was never going to constitute proof of it.

In the first week of September 2026 the same model was reported as scoring 62.7 per cent and 99.9 per cent on the same benchmark. Both numbers are correct. Both were published by the same organisation, on the same day, about the same system. The gap between them is almost entirely a fact about how the test was administered, and it is the most instructive thing to come out of the launch.

What ARC-AGI-3 asks a system to do

ARC-AGI-3, announced in March 2026, is the third generation of a benchmark series built around a specific argument: that most AI evaluations measure skill at a task, and skill at a task can be bought with enough training data. What the series tries to measure instead is how efficiently a system picks up something it has never been trained on.

The third version is the first that is interactive. A system is dropped into a small turn-based environment with no instructions and has to work out the goal, the controls and the rules by acting inside it. Every environment is calibrated on people before release, and humans solve 100 per cent of them — a design requirement rather than a result, and the detail that makes a low machine score meaningful. We set out that calibration in full in our earlier report on what ARC-AGI-3 measures and what it does not.

The two harnesses

A harness is the scaffolding around the model: what it is shown each turn, what it is allowed to carry forward, how its previous thinking is stored. The ARC Prize Foundation now runs two.

  • The Standard harness is minimal and provider-neutral. It supplies everything needed to solve each game but leaves the model responsible for deciding what to preserve in its own visible notes. Its purpose is a like-for-like comparison across companies.
  • The Provider Adapter harness uses the context-management features a model's own developer built for it. For Astra that means preserving the opaque reasoning state — which the evaluators cannot see — between requests, and compacting long conversations so earlier work can be reused.

Under the Standard harness, Astra's best result was 62.7 per cent, at maximum reasoning effort, for 26,098 dollars of compute. Under the Provider Adapter, its best was 99.9 per cent, at high reasoning effort, for 18,817 dollars.

The row that settles what is doing the work

Published alongside those headline figures is a full breakdown by reasoning effort, and one line in it is worth more than the rest combined. At the lowest setting — reasoning effort switched off entirely — the two harnesses diverge like this:

  • Standard harness, no reasoning effort: 35.2 per cent.
  • Provider Adapter harness, no reasoning effort: 96.7 per cent.

With the model's deliberate reasoning turned off, the provider harness still returns a result close to its ceiling. Whatever is producing the 99.9 per cent, it is not primarily the reasoning the benchmark was built to measure. It is the machinery around the model that carries state from one move to the next.

With reasoning effort switched off, one harness scores 35.2 per cent and the other 96.7 per cent. That is a fact about the scaffolding, not the intelligence.

The Standard column has a second oddity worth noting without over-reading: it is not monotonic. The "low" setting returns 17.5 per cent, below the 35.2 per cent recorded with no reasoning effort at all. Benchmarks at this scale carry real run-to-run variation, and a single non-monotonic step is a reminder that individual cells in these tables are noisier than their decimal places suggest.

There is also a genuinely counterintuitive result in the cost column: higher reasoning effort is cheaper, not dearer. Maximum effort under the Standard harness cost 26,098 dollars; no effort cost 49,791. The model that thinks harder solves games in fewer moves, and fewer moves means fewer calls.

How the two numbers were reported

OpenAI's announcement prints a single ARC-AGI-3 row. Astra appears at 99.9 per cent, next to 30.2 per cent for Claude Opus 5 and 7.8 per cent for GPT-5.6 Sol. A footnote records that Astra was run with OpenAI's own responses API harness, and states that the settings changed "do not specifically target ARC-AGI-3". The two comparison models were not run that way.

The disclosure is real and was made by the company itself, which is more than many launch tables offer. The difficulty is placement: a reader taking the row at face value sees a thirteenfold gap over the nearest rival, and the information needed to interpret that gap is in a footnote. The ARC Prize Foundation has since said it will report Standard and Provider Adapter results separately on its leaderboard, with each condition labelled.

What the model genuinely did achieve

Discounting the headline number is not the same as discounting the result, and the underlying achievement is substantial. Before launching ARC-AGI-3 the foundation tested roughly 500 members of the general public, not selected for puzzle-solving ability, to establish a baseline for action efficiency — the number of moves a solver needs to work an environment out. Under the provider harness, Astra used fewer actions than the median tested human on 96.0 per cent of the levels it completed, and 51.7 per cent fewer actions per level on average.

The foundation calls that a material milestone, and it is. It is also a specific claim about efficiency in moves, not about ability in general and not about cost. On cost the comparison runs the other way by several orders of magnitude: human participants were paid 115 dollars for a ninety-minute session plus 5 dollars per completed game, working out at roughly 12.78 dollars per game attempted, against the 17,000 to 26,000 dollars a full model run consumed.

The evaluators also noted a caveat about testing conditions that anyone who works with standardised assessment will recognise. In a separate, more permissive harness the model was given a sandbox and wrote its own tools — a maze solver, a combat model, a patrol tracker. The foundation observed that its human participants "did not have a code interpreter, scratch pad, etc.", so those results describe the model and its tools together rather than the model alone.

Why this is a testing story, not just an AI story

Two scores thirty-seven points apart, differing only in how the test was administered, is a problem cognitive assessment solved a long time ago and solved strictly. Standardised administration — the same instructions, the same time limit, the same permitted materials — is not an administrative courtesy wrapped around the instrument. It is part of the instrument. A test given under different conditions is a different test, and its norms no longer apply.

That is why a proctored assessment specifies what a candidate may bring into the room, why an untimed administration of a timed test produces a score that cannot be read against the published tables, and why the margin of error around any single figure is reported rather than hidden — the argument we make in why a score is a range and not a number. Machine evaluation is arriving at the same conclusions from first principles, in public, at speed.

Your own number

Where would your own score land?

The transferable habit is small and works everywhere: before believing a score, ask what it was produced under. Our own test reports a result as a position in a reference population, under stated conditions, with a range around it — the three things a bare percentage cannot supply.

Find your IQ score now!
Secure & encryptedInstant results10–20 minutes

One last thing deserves quoting directly, because it came from the organisation with the most to gain from the opposite claim. Asked whether the result represents artificial general intelligence, the ARC Prize Foundation wrote that while it believes the system represents meaningful progress towards generalisation, it is "not claiming that it is AGI" — and noted that it had said at launch that saturating the benchmark would not constitute proof of achieving it. When the benchmark's own authors decline the headline, that is worth more than the headline.

Common questions

What is ARC-AGI-3?

The third generation of the Abstraction and Reasoning Corpus benchmark, announced in March 2026 and the first that is interactive. A system is placed in a small turn-based environment with no instructions and must work out the goal, controls and rules by acting. Every environment is calibrated on human participants before release, and humans solve 100 per cent of them.

Why did GPT-6 Astra score 62.7 per cent and 99.9 per cent?

Different test harnesses. The 62.7 per cent came from the ARC Prize Foundation's provider-neutral harness, which leaves the model to decide what to record in its own visible notes. The 99.9 per cent came from a harness supplied by OpenAI that preserves the model's internal reasoning state between moves and compacts long conversations so earlier work can be reused. The model and the benchmark were the same in both cases.

Which ARC-AGI-3 number should be quoted?

For comparing models against each other, the Standard harness figure, because every system is run the same way. For describing what a system can do when used as its developer intends, the Provider Adapter figure. Quoting one model's adapter score against another model's standard score is not a comparison. The ARC Prize Foundation now labels both conditions separately on its leaderboard.

Does saturating ARC-AGI-3 mean AGI has arrived?

The organisation that built the benchmark says not. It has stated that it is not claiming the result is AGI, and that it said at the benchmark's launch that saturating it would not represent proof of achieving AGI. It defines AGI as the ability to acquire any skill a human can, as efficiently as a human can.

Sources for this story

  1. OpenAI's GPT-6 Astra on ARC-AGI-3, including the harness comparison, the per-effort score and cost table, and the human baseline — ARC Prize Foundation
  2. GPT-6 Astra ARC-AGI results page, verified scores by reasoning level and harness — ARC Prize Foundation
  3. GPT-6 Astra: A new generation of intelligence, abstract-reasoning table and footnote 1 — OpenAI
  4. Chollet, "On the Measure of Intelligence" (2019), the argument the benchmark series is built on — arXiv
  5. Standards for Educational and Psychological Testing, on standardised administration — American Educational Research Association, American Psychological Association and National Council on Measurement in Education

Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.

Share this story

Know someone who keeps seeing this number quoted without the scale it was measured on? Send it to them — it takes one tap.

Filed under#ai benchmarks#fluid reasoning#test conditions#study quality

Read the research.
Then find your own number.

Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.

Start IQ Test
Secure & encryptedInstant results10–20 minutes