OpenAI launched GPT-6 Astra on 3 September with that phrase. The one composite intelligence measure in the announcement's own comparison table puts it in fourth place.
On 3 September 2026 OpenAI announced GPT-6 Astra and described it, in the first line of the announcement, as "the world's most intelligent and aligned model". The claim travelled everywhere in the following forty-eight hours. It is worth pausing on, not because it is obviously wrong, but because the evidence that would settle it is printed further down the same page — and it does not say what the headline says.
The announcement carries a long comparison table setting Astra against GPT-5.6 Sol, Claude Fable 5.1, Claude Fable 5, Claude Opus 5 and Gemini 3.8 Flash across roughly thirty evaluations. Most of those are single-skill tests: a coding benchmark, a science question set, a cybersecurity exercise. Exactly one row is a composite measure of general capability — the Artificial Analysis Intelligence Index, an independent aggregate that blends reasoning, coding, mathematics and knowledge evaluations into one figure.
That row reads as follows.
Fourth of six, on the only line in the table that attempts to summarise general capability in a single number, published by the company making the claim. Artificial Analysis, which maintains the index independently, separately ranked Claude Fable 5.1 first of the 192 models it tracks.
A second row points the same way. On Humanity's Last Exam — a set of expert-written questions across dozens of fields, run with tools — the table gives Astra 57.2 per cent against 65.0 for Fable 5.1, 63.8 for Fable 5 and 63.6 for Opus 5. Astra is last of the four models reported. This one is worth more than the others because Anthropic publishes identical figures for those three Claude models on the same evaluation, so the comparison is not contaminated by the two labs running the test differently.
The claim is not false. It is unfalsifiable, which is a different and more interesting problem.
It would be easy, and wrong, to file this as marketing caught out. Astra leads the comparison table on most of its individual rows, and several of those leads are large. It reports 96.0 per cent on GPQA Diamond, a set of graduate-level science questions; 97.6 per cent on FrontierMath Tier 4; 57.9 per cent on Terminal-Bench 4.0 against 55.8 for Fable 5.1; 100 per cent on ExploitBench, up from 78.5 for its predecessor; and 100 per cent on long-context retrieval at 256K to 512K tokens. It also helped establish two new results on the spacing of prime numbers, improving a bound that had stood for more than eighty years.
So there is a defensible reading of "most intelligent" on which the claim holds, and a defensible reading on which it does not, and nothing in the announcement says which reading is meant. That is the actual finding here. The sentence cannot be checked, because the scale is missing.
The reason the Intelligence Index row deserves more weight than the thirty rows around it is the same reason a Full-Scale IQ exists at all. Any single test measures a narrow thing. A composite built from several different tests measures something broader, and if the underlying abilities are related, the composite carries real information that no single subtest does.
In people, they are related. Performance across vocabulary, spatial reasoning, working memory and arithmetic correlates positively — the pattern known as the general factor, which we set out in our explainer on what the g factor actually is. That correlation is what licenses a single summary number, and it is why an IQ generalises beyond the specific items on the paper.
Machine abilities do not hang together like that. A system can answer graduate physics at 96 per cent and be stopped almost completely by a coloured-grid puzzle that untrained adults solve reliably. Because the abilities are not correlated, any composite is a weighting decision — a statement about which capabilities the index author thinks matter — rather than a measurement of an underlying quantity. Change the weights and the ranking changes.
That is a genuine limitation of the index, and it cuts both ways. It is a reason not to treat 65.7 versus 61.2 as a verdict on which system is cleverer. It is also the reason "most intelligent" cannot be rescued by pointing at individual wins: if there is no single quantity being measured, there is no single ranking to top.
Everything above is one rule applied to a new subject. A number becomes a measurement when three things are stated: the scale it sits on, the population it is compared against, and the conditions under which it was produced. Strip any of the three and what is left is a figure with a decimal point in it.
None of this is a reason to dismiss the model or the benchmarks. Frontier evaluation in 2026 is a serious enterprise producing genuinely informative numbers, and the Artificial Analysis index is a reasonable attempt at the hardest part of it. The point is narrower: the superlative on the front of the announcement is doing none of that work, and the reader who wants to know what changed on 3 September has to scroll past it to the table.
The habit that protects you here is the same one that protects you from a bad online IQ result: ask which scale, which population, which conditions. Our own test reports a score as a position in a reference population with a range around it, because that is the only form in which a number of this kind means anything.
Find your IQ score now! →The useful summary of the launch is unglamorous and entirely supportable from the published tables. GPT-6 Astra is the strongest system yet released on a long list of specific, valuable capabilities, and on the one general-capability composite in its own announcement it is fourth. Both halves of that sentence are printed by OpenAI. Only one of them travelled.
There is no agreed answer, because there is no agreed scale. On the Artificial Analysis Intelligence Index, a composite of reasoning, coding, mathematics and knowledge evaluations, Claude Fable 5.1 leads at 65.7, ahead of Claude Opus 5 at 63.1 and GPT-6 Astra at 61.2 — figures published in OpenAI's own comparison table on 3 September 2026. On many individual benchmarks in that same table, Astra leads. Which model is "most intelligent" depends entirely on which measure is chosen.
No. It leads most rows, including GPQA Diamond at 96.0 per cent, FrontierMath Tier 4 at 97.6 per cent and Terminal-Bench 4.0 at 57.9 per cent. It is fourth of six on the Artificial Analysis Intelligence Index and last of the four models reported on Humanity's Last Exam with tools, at 57.2 per cent against 65.0 for Claude Fable 5.1.
An independently maintained composite score that aggregates several standardised evaluations — graduate-level science questions, broad academic knowledge, olympiad mathematics, frontier scientific knowledge and programming — into a single figure. It is the closest published analogue to a general-ability composite for machines, which is why it carries more weight than any single benchmark row.
Not as stated. A measurement requires a scale, a reference population and defined conditions. A superlative with none of the three cannot be checked either way. The same objection applies to an IQ figure quoted without its standard deviation or its norm group.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
IQ scales are defined by a human reference sample, and a language model is not in it. Placing a machine on that scale is not a hard measurement problem — it is a category error.
Frontier systems now clear reasoning benchmarks that defeated them a year ago. Put them in front of puzzles screened so that ordinary people solve every one, and the best models finish under one per cent.

The idea that everyone leans on one dominant hemisphere — logical on the left, creative on the right — traces back to a real, Nobel Prize-winning discovery. It is not what that discovery found.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →