IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment
AI & Machine IntelligenceEmerging evidence

Can AI Pass an IQ Test? On ARC-AGI-3, Humans Score 100% and AI 0.51%

Frontier systems now clear reasoning benchmarks that defeated them a year ago. Put them in front of puzzles screened so that ordinary people solve every one, and the best models finish under one per cent.

Illustration generated for IQ Metrics. No photograph is used. IQ Metrics

What you need to know

  • The ARC Prize Foundation announced ARC-AGI-3 on 25 March 2026 with one comparison on the front of the announcement: humans score 100 per cent, frontier AI scores 0.51 per cent.
  • That gap is not explained by the tasks being hard for everyone. Nothing enters an ARC-AGI evaluation set until people have solved it — ARC-AGI-2 was calibrated on more than 400 untrained participants, and a task was admitted only if at least two of them solved it within two attempts.
  • A benchmark percentage is not an IQ. 0.51 per cent is the share of tasks completed; an IQ is a position inside a reference population with a stated mean and standard deviation. Neither number converts into the other, and no norms for machines exist.
  • What the two formats do share is novelty. A matrix item on a human test and an ARC grid both ask for a rule you have never met, inferred from two or three worked examples — which is why the result is worth reading even though it is not a score on any human scale.

Most months now bring a report of a model clearing some benchmark that was supposed to hold for years. Set against that steady drumbeat, a figure published in March 2026 is worth sitting with. On ARC-AGI-3, the newest benchmark from the ARC Prize Foundation, the organisation's own summary of the developer preview reads: humans score 100 per cent, frontier AI scores 0.51 per cent. Both halves of that sentence carry weight, and the first half carries more of it than it looks.

A benchmark built to be unstudyable

The series began with an argument rather than a leaderboard. In a 2019 paper called "On the Measure of Intelligence", François Chollet made the case that nearly every AI benchmark of the period measured skill at a task, and that skill at a task can simply be bought with enough data and enough training. What he wanted a benchmark to measure instead was how efficiently a system picks up a skill it has never been trained on. The corpus he published alongside the paper — the Abstraction and Reasoning Corpus, released in November 2019 — held 1,000 tasks split into 400 training, 400 public evaluation and 200 held-back test items.

The format is disarmingly plain. You are shown a small coloured grid and the grid it turns into, two or three times over. Then you are shown a new input grid and asked to produce the output. Nobody tells you the rule. Every task has a different one — reflect the shape, fill the enclosed region, count the odd colour out, continue the sequence to the edge — and none of them carries over to the next task. There is nothing to revise for, which is the entire design goal.

ARC-AGI-2 followed on 24 March 2025, harder in construction and, more importantly for reading the headline number, calibrated on people. Tasks were tested on more than 400 untrained participants working individually in supervised sessions, and a task was allowed into the evaluation set only once at least two of them had solved it in two attempts or fewer. That screening is what turns "AI scores low" into a claim worth making at all.

  • It removes the easiest explanation. A benchmark nobody can do tells you nothing about the gap between people and machines; a benchmark everybody can do tells you a great deal.
  • It makes the human figure a property of the task set rather than a marketing number. The 100 per cent is a completion bar every admitted item had to clear, not an average over some sample of especially clever volunteers.
  • It keeps the difficulty in the right place. Tasks are meant to be novel, not obscure — no specialist knowledge, no vocabulary, no arithmetic beyond counting.
  • And it makes the benchmark checkable. The tasks are public, the calibration procedure is documented, and anyone can sit down and try the puzzles themselves.

What the third version changed

ARC-AGI-3, announced on 25 March 2026, is the first version in the series that is interactive. Rather than a static grid mapped to another grid, a system is dropped into a small turn-based environment — a game world with no instructions — and has to work out the goal, the controls and the rules by acting inside it. The information is not handed over up front. It has to be gone and got.

That is a much larger ask than it sounds, and it is the reason the numbers separate so violently. A static ARC task gives you everything at once and asks for one inference. An interactive environment asks for a loop: form a guess about what the world does, take an action that would distinguish that guess from a rival one, read the result, revise, and keep a running model of a system nobody described to you. People do this without noticing. It is roughly what a child does with an unfamiliar toy in the first thirty seconds.

A benchmark that every human solves and almost no machine solves is not measuring difficulty. It is measuring one specific kind of learning.

Why 0.51 per cent is not an IQ of anything

The temptation, once a machine and a human are compared on the same tasks, is to reach for the familiar scale and say the machine has an IQ of something. It does not, and the reason is worth stating precisely, because the same confusion turns up whenever a model is fed a test paper. A benchmark score is a raw percentage: tasks completed out of tasks attempted. An IQ is not a percentage of anything. It is a position within a reference population, expressed on a scale with a deliberately chosen mean — almost always 100 — and a deliberately chosen standard deviation, which is the measure of how far apart people's results fall, usually 15.

To put a machine on that scale you would need a reference population the machine belonged to, and a norming study in which the test was administered to a representative sample of that population under standard conditions. None of that exists, and it is not obvious what it would even mean. We have set out the full argument separately: a model answering IQ items produces a count of correct answers, and the step from there to an IQ requires norms nobody has collected.

Do not take any specific leaderboard figure from an article, including this one, as current. The earlier versions have moved fast: ARC-AGI-1 has been largely worked through by frontier reasoning systems, and scores on ARC-AGI-2 climbed steeply through 2025 and 2026 from a starting point in the low single digits. Published rankings change from week to week and are reported under cost caps that differ between entries. The 0.51 per cent quoted here is the ARC Prize Foundation's own figure for the ARC-AGI-3 developer preview at launch; check the live leaderboard before repeating it.

What this says about the puzzles on a human test

There is a reason matrix puzzles sit near the centre of most non-verbal reasoning tests, and the ARC results illustrate it better than any textbook paragraph. A matrix item is chosen precisely because it leans as little as possible on what you happen to know — no vocabulary, no schooling, no cultural reference — and as much as possible on working out a rule from the evidence in front of you and applying it once. Our walkthrough of how to solve a matrix item by naming the rule out loud isolates the same skill the ARC grids are built around.

What the benchmark demonstrates is how much of the difficulty in these items comes from novelty alone, once everything else is stripped out. Systems that can write competent code, pass professional examinations and summarise a legal filing are stopped almost completely by puzzles with no content in them. That is not a claim that people are cleverer in general — on a very large number of tasks the comparison runs decisively the other way — but it is a clean demonstration that "performs well on tests" and "learns a new rule from two examples" are separable abilities. That separation is exactly what a good reasoning item is trying to catch, and it is why how a test is constructed matters more than how hard its questions feel.

The comparison does not become automatic on the human side either. A person sitting a reasoning test produces a raw count of correct answers, and that count only becomes interpretable once it is placed against a reference sample — which is what the IQ percentile calculator does, converting a score on a stated scale into the share of the population sitting below it. A benchmark percentage has no equivalent step available, because there is no population of machines for a model to be a member of. That is not a technicality about scoring; it is the whole difference between a measurement and a tally.

Your own number

Where would your own score land?

The quickest way to understand what a fluid-reasoning item is doing is to sit one. Our IQ test is built around exactly this kind of rule-finding, and the percentile calculator will tell you where a score stands in a reference population — the step a raw benchmark percentage does not have and cannot fake.

Find your IQ score now!
Secure & encryptedInstant results10–20 minutes

How to read the next headline about this

  • Ask whether the tasks were calibrated on people. If nobody checked that humans can do them, "AI fails" is not informative.
  • Ask what the number is a percentage of. Tasks solved, questions correct and items completed within a cost budget are three different quantities.
  • Ask whether the benchmark was public before the model was trained. A benchmark sitting in the training data measures memory, not reasoning.
  • Ask who is quoting it, and which version. Results on ARC-AGI-1, ARC-AGI-2 and ARC-AGI-3 are not comparable to one another, and benchmark scores are marketing as often as they are measurement.
  • And keep the scale question in front of you: a percentage on a benchmark and a score on a normed test are different kinds of number, however tempting the comparison.

None of this makes the 0.51 per cent a permanent state of affairs. The whole series exists because each previous version eventually stopped separating systems, and the sensible expectation is that this one will too. The durable point is narrower and more useful: the thing these puzzles isolate — picking up an unfamiliar rule quickly, from very little evidence — is a real and separable component of what cognitive tests try to measure, and it has proved the hardest to buy with scale. It is also the component the general factor in human testing leans on most heavily, and the one most easily outsourced when a machine is close at hand.

Common questions

Can AI pass an IQ test?

A model can be shown IQ-style items and can answer many of them correctly, but that produces a count of correct answers, not an IQ. An IQ is a position within a reference population, defined by a mean and a standard deviation established in a norming study. No such norming exists for machines, so the score cannot be placed on the scale. On ARC-AGI-3, a benchmark of novel reasoning puzzles announced in March 2026, the ARC Prize Foundation reported humans at 100 per cent and frontier AI at 0.51 per cent.

What is ARC-AGI-3?

It is the third version of the Abstraction and Reasoning Corpus benchmark, announced on 25 March 2026, and the first that is interactive. Instead of mapping one grid to another, a system is placed in a small turn-based environment with no instructions and must work out the goal and the rules by acting. The series was created by François Chollet, whose 2019 paper "On the Measure of Intelligence" argued that benchmarks should measure how efficiently a system acquires a new skill rather than how well it performs a trained one.

Why do humans score 100% on ARC-AGI?

Because that is a design requirement, not a coincidence. Candidate tasks are tested on people before release, and a task enters the evaluation set only if humans solve it. For ARC-AGI-2, more than 400 untrained participants were tested and a task was admitted only when at least two of them solved it in two attempts or fewer. The screening is what makes a low machine score meaningful — it rules out the possibility that the tasks are simply impossible.

Does an AI benchmark score tell you anything about human intelligence?

Indirectly, and mostly about test design. It shows how much of the difficulty in a good reasoning item comes from novelty alone, once knowledge, vocabulary and training are stripped out. It says nothing about any individual person's ability, and a benchmark percentage cannot be converted into an IQ or a percentile.

Sources for this story

  1. ARC-AGI-3 announcement and the reported human and frontier-model figures — ARC Prize Foundation
  2. ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence, technical report (2026) — arXiv
  3. ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems (2025), including the human calibration procedure — arXiv
  4. Chollet, "On the Measure of Intelligence" (2019) — arXiv
  5. ARC Prize 2025: Technical Report — ARC Prize Foundation
  6. Standards for Educational and Psychological Testing, on norming and the meaning of a standard score — American Educational Research Association, American Psychological Association and National Council on Measurement in Education

Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.

Share this story

Know someone who keeps seeing this number quoted without the scale it was measured on? Send it to them — it takes one tap.

Filed under#ai benchmarks#fluid reasoning#raven matrices#study quality#scores and scales

Read the research.
Then find your own number.

Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.

Start IQ Test
Secure & encryptedInstant results10–20 minutes