Frontier systems now clear reasoning benchmarks that defeated them a year ago. Put them in front of puzzles screened so that ordinary people solve every one, and the best models finish under one per cent.
Most months now bring a report of a model clearing some benchmark that was supposed to hold for years. Set against that steady drumbeat, a figure published in March 2026 is worth sitting with. On ARC-AGI-3, the newest benchmark from the ARC Prize Foundation, the organisation's own summary of the developer preview reads: humans score 100 per cent, frontier AI scores 0.51 per cent. Both halves of that sentence carry weight, and the first half carries more of it than it looks.
The series began with an argument rather than a leaderboard. In a 2019 paper called "On the Measure of Intelligence", François Chollet made the case that nearly every AI benchmark of the period measured skill at a task, and that skill at a task can simply be bought with enough data and enough training. What he wanted a benchmark to measure instead was how efficiently a system picks up a skill it has never been trained on. The corpus he published alongside the paper — the Abstraction and Reasoning Corpus, released in November 2019 — held 1,000 tasks split into 400 training, 400 public evaluation and 200 held-back test items.
The format is disarmingly plain. You are shown a small coloured grid and the grid it turns into, two or three times over. Then you are shown a new input grid and asked to produce the output. Nobody tells you the rule. Every task has a different one — reflect the shape, fill the enclosed region, count the odd colour out, continue the sequence to the edge — and none of them carries over to the next task. There is nothing to revise for, which is the entire design goal.
ARC-AGI-2 followed on 24 March 2025, harder in construction and, more importantly for reading the headline number, calibrated on people. Tasks were tested on more than 400 untrained participants working individually in supervised sessions, and a task was allowed into the evaluation set only once at least two of them had solved it in two attempts or fewer. That screening is what turns "AI scores low" into a claim worth making at all.
ARC-AGI-3, announced on 25 March 2026, is the first version in the series that is interactive. Rather than a static grid mapped to another grid, a system is dropped into a small turn-based environment — a game world with no instructions — and has to work out the goal, the controls and the rules by acting inside it. The information is not handed over up front. It has to be gone and got.
That is a much larger ask than it sounds, and it is the reason the numbers separate so violently. A static ARC task gives you everything at once and asks for one inference. An interactive environment asks for a loop: form a guess about what the world does, take an action that would distinguish that guess from a rival one, read the result, revise, and keep a running model of a system nobody described to you. People do this without noticing. It is roughly what a child does with an unfamiliar toy in the first thirty seconds.
A benchmark that every human solves and almost no machine solves is not measuring difficulty. It is measuring one specific kind of learning.
The temptation, once a machine and a human are compared on the same tasks, is to reach for the familiar scale and say the machine has an IQ of something. It does not, and the reason is worth stating precisely, because the same confusion turns up whenever a model is fed a test paper. A benchmark score is a raw percentage: tasks completed out of tasks attempted. An IQ is not a percentage of anything. It is a position within a reference population, expressed on a scale with a deliberately chosen mean — almost always 100 — and a deliberately chosen standard deviation, which is the measure of how far apart people's results fall, usually 15.
To put a machine on that scale you would need a reference population the machine belonged to, and a norming study in which the test was administered to a representative sample of that population under standard conditions. None of that exists, and it is not obvious what it would even mean. We have set out the full argument separately: a model answering IQ items produces a count of correct answers, and the step from there to an IQ requires norms nobody has collected.
There is a reason matrix puzzles sit near the centre of most non-verbal reasoning tests, and the ARC results illustrate it better than any textbook paragraph. A matrix item is chosen precisely because it leans as little as possible on what you happen to know — no vocabulary, no schooling, no cultural reference — and as much as possible on working out a rule from the evidence in front of you and applying it once. Our walkthrough of how to solve a matrix item by naming the rule out loud isolates the same skill the ARC grids are built around.
What the benchmark demonstrates is how much of the difficulty in these items comes from novelty alone, once everything else is stripped out. Systems that can write competent code, pass professional examinations and summarise a legal filing are stopped almost completely by puzzles with no content in them. That is not a claim that people are cleverer in general — on a very large number of tasks the comparison runs decisively the other way — but it is a clean demonstration that "performs well on tests" and "learns a new rule from two examples" are separable abilities. That separation is exactly what a good reasoning item is trying to catch, and it is why how a test is constructed matters more than how hard its questions feel.
The comparison does not become automatic on the human side either. A person sitting a reasoning test produces a raw count of correct answers, and that count only becomes interpretable once it is placed against a reference sample — which is what the IQ percentile calculator does, converting a score on a stated scale into the share of the population sitting below it. A benchmark percentage has no equivalent step available, because there is no population of machines for a model to be a member of. That is not a technicality about scoring; it is the whole difference between a measurement and a tally.
The quickest way to understand what a fluid-reasoning item is doing is to sit one. Our IQ test is built around exactly this kind of rule-finding, and the percentile calculator will tell you where a score stands in a reference population — the step a raw benchmark percentage does not have and cannot fake.
Find your IQ score now! →None of this makes the 0.51 per cent a permanent state of affairs. The whole series exists because each previous version eventually stopped separating systems, and the sensible expectation is that this one will too. The durable point is narrower and more useful: the thing these puzzles isolate — picking up an unfamiliar rule quickly, from very little evidence — is a real and separable component of what cognitive tests try to measure, and it has proved the hardest to buy with scale. It is also the component the general factor in human testing leans on most heavily, and the one most easily outsourced when a machine is close at hand.
A model can be shown IQ-style items and can answer many of them correctly, but that produces a count of correct answers, not an IQ. An IQ is a position within a reference population, defined by a mean and a standard deviation established in a norming study. No such norming exists for machines, so the score cannot be placed on the scale. On ARC-AGI-3, a benchmark of novel reasoning puzzles announced in March 2026, the ARC Prize Foundation reported humans at 100 per cent and frontier AI at 0.51 per cent.
It is the third version of the Abstraction and Reasoning Corpus benchmark, announced on 25 March 2026, and the first that is interactive. Instead of mapping one grid to another, a system is placed in a small turn-based environment with no instructions and must work out the goal and the rules by acting. The series was created by François Chollet, whose 2019 paper "On the Measure of Intelligence" argued that benchmarks should measure how efficiently a system acquires a new skill rather than how well it performs a trained one.
Because that is a design requirement, not a coincidence. Candidate tasks are tested on people before release, and a task enters the evaluation set only if humans solve it. For ARC-AGI-2, more than 400 untrained participants were tested and a task was admitted only when at least two of them solved it in two attempts or fewer. The screening is what makes a low machine score meaningful — it rules out the possibility that the tasks are simply impossible.
Indirectly, and mostly about test design. It shows how much of the difficulty in a good reasoning item comes from novelty alone, once knowledge, vocabulary and training are stripped out. It says nothing about any individual person's ability, and a benchmark percentage cannot be converted into an IQ or a percentile.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
IQ scales are defined by a human reference sample, and a language model is not in it. Placing a machine on that scale is not a hard measurement problem — it is a category error.
Three rows, three columns, two things changing at once. The worked solution is below — and so is an honest account of what solving it does and does not tell you about yourself.
GPT-6 Astra scored 62.7 per cent and 99.9 per cent on the same benchmark in the same week. The difference is not the model. It is how the test was administered.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →