IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

IQ Metrics

Is AI Smarter Than Humans?

Understanding IQ

Is AI Smarter Than Humans? What the 2026 Benchmarks Show

Ask whether AI is smarter than humans and the honest answer depends entirely on the task. Frontier systems now beat expert humans on graduate-level science questions and lose to ordinary people on puzzles built to be unfamiliar. Here is what the 2026 benchmarks actually measure, and why one score never settles it.

Chart comparing machine and human performance across four task types, showing machines ahead on graduate science questions and coding, level on novel puzzle solving, and behind on tasks never seen before

Is AI smarter than humans? On narrow, well-specified tasks the answer in 2026 is often yes, and increasingly by a wide margin. On problems a system has never seen before, described by nobody, with no worked example to copy, people still hold the advantage. The reason both statements are true at once is that "smarter" is not one axis, and the tests that make machines look unbeatable and the tests that stop them cold are measuring genuinely different things.

This article takes the comparison domain by domain, using figures published by the labs and by independent evaluators in 2026, and then explains the structural difference that makes a single answer impossible.

What AI already does better than most people

The clearest machine wins are on tasks with a correct answer, a large body of prior examples, and no requirement to act in the world. On GPQA Diamond, a set of graduate-level biology, chemistry and physics questions written to be hard for people with access to a search engine, OpenAI reported GPT-6 Astra at 96.0 percent in September 2026. Gemini 3.1 Pro sits at 94.3 percent. Both are above the performance of domain experts answering outside their own speciality.

The same pattern holds across coding, long-document retrieval and structured professional work. These are not trick results. They are real, they are reproducible, and they describe abilities that took people years of training to acquire.

  • Recall and synthesis at volume. No person holds the contents of a technical literature in working memory. A model effectively does.
  • Speed. Work that takes an expert a day is returned in minutes, which changes what is worth attempting.
  • Consistency. A model does not get tired on the four-hundredth item, which is exactly where human scorers drift.
  • Breadth of surface knowledge. Competence across far more fields than any individual can maintain.

Where humans still hold the edge

The sharpest counterexample is ARC-AGI-3, a benchmark from the ARC Prize Foundation built specifically to test learning rather than recall. A system is dropped into a small turn-based environment with no instructions and has to work out the goal, the controls and the rules by acting inside it. Crucially, every environment is calibrated on people first: humans solve 100 percent of them, because a task is only admitted once people have shown it can be done.

That calibration is what makes the comparison meaningful, and it is the detail most coverage drops. We set out the full argument in our report on what ARC-AGI-3 measures. The short version: when the novelty is real and the instructions are absent, the gap between an ordinary adult and a frontier system has been enormous.

That gap is now closing fast, and the way it closed is instructive. In September 2026 the ARC Prize Foundation reported GPT-6 Astra at 62.7 percent on its neutral, provider-independent test harness — and 99.9 percent on a harness supplied by the model’s own developer, which preserves the system’s reasoning state between moves. Same model, same benchmark, same week. The scaffolding around the model accounted for most of the difference.

Is AI smarter than humans at learning new things?

This is the question the benchmark was built to ask, and the 2026 answer is genuinely mixed. On action efficiency — how many moves a solver needs to work an unfamiliar environment out — the ARC Prize Foundation found that Astra used fewer actions than the median tested human on 96 percent of the levels it completed, and 51.7 percent fewer actions per level on average. By that measure the machine matched and passed human parity.

Cost tells a different story. The human baseline came from around 500 members of the public, paid roughly 12.78 dollars per game attempted. The model runs that produced those scores cost between 17,332 and 26,098 dollars. The system that learns as efficiently as a person in moves does so at several thousand times the price in resources.

Both numbers are real, and neither alone answers the headline question. That is the pattern to expect from here.

Chart comparing machine and human performance across four task types, showing machines ahead on graduate science questions and coding, level on novel puzzle solving, and behind on tasks never seen before
Chart comparing machine and human performance across four task types, showing machines ahead on graduate science questions and coding, level on novel puzzle solving, and behind on tasks never seen before
Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now!
Secure & encryptedInstant results10–20 minutes

Why smarter is the wrong question

In people, mental abilities correlate. Someone who scores well on vocabulary tends to score well on spatial reasoning and on arithmetic, and the pattern is consistent enough that a single summary number carries real information. That pattern is the reason IQ works at all — we explain the underlying statistics in our explainer on the g factor.

Machines do not show that structure. A system can answer graduate physics correctly and then fail a coloured-grid puzzle that a ten-year-old solves in thirty seconds. Its abilities do not hang together, so no single number summarises them, and any ranking against a person depends entirely on which task you picked. This is not a temporary measurement problem. It is a real difference in how the two kinds of system are built.

  • Human ability is correlated; one score generalises across many tasks.
  • Machine ability is jagged; world-class in one domain, below average in the next, with no reliable pattern.
  • So a comparison needs a task, and any claim without one is not a measurement.

Does a benchmark score mean the same as an IQ score?

No, and the distinction matters more than it sounds. A benchmark result is a raw percentage: items solved out of items attempted. An IQ is not a percentage of anything. It is a position within a reference population, expressed on a scale with a defined mean, almost always 100, and a defined standard deviation, usually 15.

Turning a raw count into that position requires a norming study: the same test, administered under standard conditions, to a representative sample of the population the score will be read against. No such sample exists for machines, and it is not clear what one would even be. So a model can score 96 percent on a science exam without that number converting into any IQ at all. We take that argument apart properly in the article on what ChatGPT’s IQ score really is.

The same logic governs human scores, which is why a raw count on a reasoning test is meaningless until it is placed against a reference sample. That placement is exactly what the IQ percentile calculator does, and it is the step a benchmark percentage has no equivalent for.

When will AI be smarter than humans overall?

Predictions from serious people vary by decades, which is itself the most useful fact about them. Geoffrey Hinton has said he expects machines to surpass human intelligence within about twenty years. Others working on the same systems put it sooner or reject the framing entirely. There is no measurement that would settle the disagreement, because there is no agreed test — the problem we work through in the explainer on what AGI actually means.

What can be said with confidence is narrower. The hardest evaluations still defeat the best systems: on Humanity’s Last Exam, a set of expert-written questions across dozens of fields, the strongest reported result in September 2026 was 65.0 percent, meaning better than a third of the questions remained unanswered. Benchmarks that were supposed to hold for years keep falling, and new ones keep being built because the old ones stop separating anything.

What this means for measuring your own intelligence

None of this changes what a cognitive test does for a person. The reason matrix puzzles sit near the centre of most non-verbal reasoning tests is that they lean as little as possible on what you happen to know and as much as possible on working out a rule from the evidence in front of you. That the same format is what machines found hardest is a point in favour of the format, not against it.

If the comparison has made you curious about your own reasoning rather than a model’s, our IQ test is built around exactly this kind of rule-finding, and the result is reported the way a score has to be reported to mean anything: as a position in a reference population, with a range around it. For what those numbers do and do not predict, the evidence on outcomes is a better guide than any headline about machines.

The durable conclusion is unglamorous. Machines are now better than most people at a growing list of specific things, worse at a shrinking list, and not comparable at all on the single scale the question implies. Anyone offering you one number for it is selling something.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.