IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment
Myths & Fact ChecksWell established

The "World's Most Intelligent Model" Ranks Fourth on Its Own Chart

OpenAI launched GPT-6 Astra on 3 September with that phrase. The one composite intelligence measure in the announcement's own comparison table puts it in fourth place.

Illustration generated for IQ Metrics. No photograph is used. IQ Metrics

What you need to know

  • OpenAI introduced GPT-6 Astra on 3 September 2026 as "the world's most intelligent and aligned model". On the Artificial Analysis Intelligence Index printed in OpenAI's own comparison table, Astra scores 61.2 and sits fourth, behind Claude Fable 5.1 at 65.7, Claude Opus 5 at 63.1 and Claude Fable 5 at 62.1.
  • On Humanity's Last Exam with tools, the same table reports Astra at 57.2 per cent against 65.0 for Fable 5.1. Both companies publish identical figures for the three Claude models on that evaluation, so this particular comparison is like-for-like rather than a methodology dispute.
  • None of this makes the launch claim dishonest. Astra leads its comparison table on most individual rows — graduate science, terminal tasks, long-context retrieval, exploit development. The claim is unfalsifiable rather than false, because "most intelligent" has no stated scale.
  • That is the same defect this section flags in human score reporting. A number is only a measurement once someone says which scale it sits on and which population it is being read against.

On 3 September 2026 OpenAI announced GPT-6 Astra and described it, in the first line of the announcement, as "the world's most intelligent and aligned model". The claim travelled everywhere in the following forty-eight hours. It is worth pausing on, not because it is obviously wrong, but because the evidence that would settle it is printed further down the same page — and it does not say what the headline says.

What the announcement's own table shows

The announcement carries a long comparison table setting Astra against GPT-5.6 Sol, Claude Fable 5.1, Claude Fable 5, Claude Opus 5 and Gemini 3.8 Flash across roughly thirty evaluations. Most of those are single-skill tests: a coding benchmark, a science question set, a cybersecurity exercise. Exactly one row is a composite measure of general capability — the Artificial Analysis Intelligence Index, an independent aggregate that blends reasoning, coding, mathematics and knowledge evaluations into one figure.

That row reads as follows.

  • Claude Fable 5.1 — 65.7
  • Claude Opus 5 — 63.1
  • Claude Fable 5 — 62.1
  • GPT-6 Astra — 61.2
  • GPT-5.6 Sol — 60.9
  • Gemini 3.8 Flash — 58.7

Fourth of six, on the only line in the table that attempts to summarise general capability in a single number, published by the company making the claim. Artificial Analysis, which maintains the index independently, separately ranked Claude Fable 5.1 first of the 192 models it tracks.

A second row points the same way. On Humanity's Last Exam — a set of expert-written questions across dozens of fields, run with tools — the table gives Astra 57.2 per cent against 65.0 for Fable 5.1, 63.8 for Fable 5 and 63.6 for Opus 5. Astra is last of the four models reported. This one is worth more than the others because Anthropic publishes identical figures for those three Claude models on the same evaluation, so the comparison is not contaminated by the two labs running the test differently.

The claim is not false. It is unfalsifiable, which is a different and more interesting problem.

Why the claim is still not a lie

It would be easy, and wrong, to file this as marketing caught out. Astra leads the comparison table on most of its individual rows, and several of those leads are large. It reports 96.0 per cent on GPQA Diamond, a set of graduate-level science questions; 97.6 per cent on FrontierMath Tier 4; 57.9 per cent on Terminal-Bench 4.0 against 55.8 for Fable 5.1; 100 per cent on ExploitBench, up from 78.5 for its predecessor; and 100 per cent on long-context retrieval at 256K to 512K tokens. It also helped establish two new results on the spacing of prime numbers, improving a bound that had stood for more than eighty years.

So there is a defensible reading of "most intelligent" on which the claim holds, and a defensible reading on which it does not, and nothing in the announcement says which reading is meant. That is the actual finding here. The sentence cannot be checked, because the scale is missing.

One caveat applies to every figure in that table and is disclosed at the foot of it: "Evaluation scores are the maximum at any effort." Each number is the best result across reasoning-effort settings rather than a result at a fixed budget. A human score reported that way — best of several sittings, no fixed conditions — would not be accepted by any testing standard.

Why a composite index is the closest thing to a fair question

The reason the Intelligence Index row deserves more weight than the thirty rows around it is the same reason a Full-Scale IQ exists at all. Any single test measures a narrow thing. A composite built from several different tests measures something broader, and if the underlying abilities are related, the composite carries real information that no single subtest does.

In people, they are related. Performance across vocabulary, spatial reasoning, working memory and arithmetic correlates positively — the pattern known as the general factor, which we set out in our explainer on what the g factor actually is. That correlation is what licenses a single summary number, and it is why an IQ generalises beyond the specific items on the paper.

Machine abilities do not hang together like that. A system can answer graduate physics at 96 per cent and be stopped almost completely by a coloured-grid puzzle that untrained adults solve reliably. Because the abilities are not correlated, any composite is a weighting decision — a statement about which capabilities the index author thinks matter — rather than a measurement of an underlying quantity. Change the weights and the ranking changes.

That is a genuine limitation of the index, and it cuts both ways. It is a reason not to treat 65.7 versus 61.2 as a verdict on which system is cleverer. It is also the reason "most intelligent" cannot be rescued by pointing at individual wins: if there is no single quantity being measured, there is no single ranking to top.

The same rule this section applies to human scores

Everything above is one rule applied to a new subject. A number becomes a measurement when three things are stated: the scale it sits on, the population it is compared against, and the conditions under which it was produced. Strip any of the three and what is left is a figure with a decimal point in it.

  • An IQ of 130 means nothing until you know the standard deviation — it is the 97.7th percentile on a 15-point scale and the 96.9th on a 16-point scale, as we set out in our piece on why one percentile has three numbers.
  • A benchmark percentage means nothing until you know the harness, the effort setting and whether the figure is a single run or the best of several.
  • A composite means nothing until you know what went into it and how it was weighted.
  • And "most intelligent", with none of the three, means nothing at all.

None of this is a reason to dismiss the model or the benchmarks. Frontier evaluation in 2026 is a serious enterprise producing genuinely informative numbers, and the Artificial Analysis index is a reasonable attempt at the hardest part of it. The point is narrower: the superlative on the front of the announcement is doing none of that work, and the reader who wants to know what changed on 3 September has to scroll past it to the table.

Your own number

Where would your own score land?

The habit that protects you here is the same one that protects you from a bad online IQ result: ask which scale, which population, which conditions. Our own test reports a score as a position in a reference population with a range around it, because that is the only form in which a number of this kind means anything.

Find your IQ score now!
Secure & encryptedInstant results10–20 minutes

The useful summary of the launch is unglamorous and entirely supportable from the published tables. GPT-6 Astra is the strongest system yet released on a long list of specific, valuable capabilities, and on the one general-capability composite in its own announcement it is fourth. Both halves of that sentence are printed by OpenAI. Only one of them travelled.

Common questions

What is the most intelligent AI model in 2026?

There is no agreed answer, because there is no agreed scale. On the Artificial Analysis Intelligence Index, a composite of reasoning, coding, mathematics and knowledge evaluations, Claude Fable 5.1 leads at 65.7, ahead of Claude Opus 5 at 63.1 and GPT-6 Astra at 61.2 — figures published in OpenAI's own comparison table on 3 September 2026. On many individual benchmarks in that same table, Astra leads. Which model is "most intelligent" depends entirely on which measure is chosen.

Did GPT-6 Astra top every benchmark in its announcement?

No. It leads most rows, including GPQA Diamond at 96.0 per cent, FrontierMath Tier 4 at 97.6 per cent and Terminal-Bench 4.0 at 57.9 per cent. It is fourth of six on the Artificial Analysis Intelligence Index and last of the four models reported on Humanity's Last Exam with tools, at 57.2 per cent against 65.0 for Claude Fable 5.1.

What is the Artificial Analysis Intelligence Index?

An independently maintained composite score that aggregates several standardised evaluations — graduate-level science questions, broad academic knowledge, olympiad mathematics, frontier scientific knowledge and programming — into a single figure. It is the closest published analogue to a general-ability composite for machines, which is why it carries more weight than any single benchmark row.

Is "most intelligent" a measurable claim?

Not as stated. A measurement requires a scale, a reference population and defined conditions. A superlative with none of the three cannot be checked either way. The same objection applies to an IQ figure quoted without its standard deviation or its norm group.

Sources for this story

  1. GPT-6 Astra: A new generation of intelligence, including the full benchmark comparison table and its footnotes — OpenAI
  2. Introducing Claude Fable 5.1 and Claude Mythos 5.1, including Humanity's Last Exam and Terminal-Bench figures — Anthropic
  3. Artificial Analysis Intelligence Index, model leaderboard — Artificial Analysis
  4. Standards for Educational and Psychological Testing, on score interpretation and the reporting of scales — American Educational Research Association, American Psychological Association and National Council on Measurement in Education

Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.

Share this story

Know someone who keeps seeing this number quoted without the scale it was measured on? Send it to them — it takes one tap.

Filed under#ai benchmarks#myths and fact checks#scores and scales#study quality

Read the research.
Then find your own number.

Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.

Start IQ Test
Secure & encryptedInstant results10–20 minutes