IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

What Is AGI?

Understanding IQ

What Is AGI, and How Would Anyone Measure It?

Artificial general intelligence means a system that can learn and apply knowledge across the full range of tasks a person can, rather than excelling at a narrow set. The trouble is that the major labs each publish a different definition, none of them converts into a test, and so claims that AGI has arrived cannot be checked either way.

Chart comparing four published definitions of artificial general intelligence from OpenAI, Google, Microsoft and the ARC Prize Foundation, showing how little they overlap on what would count

AGI stands for artificial general intelligence: an AI system with the ability to learn, reason and apply knowledge across the full range of tasks a person can handle, rather than performing brilliantly in one narrow domain. That much is broadly agreed. Almost nothing after it is.

The organisations building these systems publish materially different definitions, none of those definitions specifies a test, and there is no reference population against which a result could be scored. So when a launch announcement is described as the arrival of AGI, there is no procedure available to check the claim. That is a measurement problem, and human intelligence testing solved a version of it a century ago.

Why nobody agrees on what AGI means

The published definitions differ on what would count, and the differences are not cosmetic.

  • OpenAI: autonomous systems that outperform humans at most economically valuable work.
  • Google: AI that is at least as capable as humans at most cognitive tasks.
  • Microsoft: the point at which an AI can match human performance at all tasks.
  • ARC Prize Foundation: a system’s ability to acquire any skill a human can, as efficiently as a human can.

Read them side by side and the gaps open up. OpenAI’s is economic and could in principle be satisfied without the system ever learning anything new. Microsoft’s says all tasks, which is a far higher bar than most. The ARC Prize definition is the only one that puts learning efficiency at the centre, which is also the only version that resembles what psychologists mean by general ability.

The disagreement is not merely academic, because each definition implies a different finish line and a different set of evidence. An economic definition is settled by labour-market data. A cognitive-task definition is settled by evaluations. A learning-efficiency definition is settled by putting a system in front of something genuinely unfamiliar and counting how much evidence it needs. A system could satisfy one and fail another on the same day, which is roughly where things stand.

A term that means four things does not support a yes-or-no answer. It supports four of them.

What is the difference between AI, AGI and superintelligence?

The three terms describe a ladder, and most confusion comes from using them interchangeably.

  • Narrow AI is everything in use today: systems that perform specific tasks, sometimes far better than people, without the ability to transfer that competence to an unrelated problem. A model that writes excellent code and cannot fold a towel is narrow, however impressive the code.
  • AGI would match human ability across the full breadth of tasks rather than in a slice of them. Breadth is the operative word: the claim is about range, not peak performance.
  • Superintelligence would exceed human ability across that same full range. It is a claim about what comes after AGI, and is entirely hypothetical.

The distinction that matters for measurement is the first one. Narrow ability is straightforward to test, because you can specify the task. Breadth is not, because you would have to specify every task — which is the problem the next section is about.

How would anyone actually measure AGI?

The most serious attempt starts from a 2019 paper by François Chollet called "On the Measure of Intelligence", which argued that nearly every AI benchmark of the period measured skill at a task — and that skill at a task can simply be bought with enough data and training. What a benchmark should measure instead, on that argument, is how efficiently a system picks up a skill it was never trained on.

That produced the ARC-AGI series, built around puzzles with no vocabulary, no cultural reference and no arithmetic beyond counting, where every task has a different rule and nothing carries over. The distinction it targets is the one human testing already draws between fluid reasoning — working out something new — and crystallized ability, which is what you have accumulated. The general factor in human ability leans most heavily on the first.

Older proposals still get cited and are worth knowing. The Turing test asked whether a machine could be told apart from a person in conversation. Marvin Minsky suggested in 1970 that a machine should be able to read Shakespeare, grease a car, play office politics and tell a joke. The coffee test, associated with Steve Wozniak, proposes that AGI arrives when a machine can walk into an unfamiliar house and make a pot of coffee. Each captures something. None produces a score.

Chart comparing four published definitions of artificial general intelligence from OpenAI, Google, Microsoft and the ARC Prize Foundation, showing how little they overlap on what would count
Chart comparing four published definitions of artificial general intelligence from OpenAI, Google, Microsoft and the ARC Prize Foundation, showing how little they overlap on what would count
Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now!

Secure & encryptedInstant results10–20 minutes

Did GPT-6 Astra just reach AGI?

In September 2026 OpenAI announced GPT-6 Astra, and reported 99.9 percent on ARC-AGI-3 — the interactive benchmark built specifically to resist this kind of result. Coverage framed it as the arrival of the AGI era.

The organisation that owns the benchmark disagrees, in writing. The ARC Prize Foundation’s own analysis states plainly that while it believes the system represents meaningful progress towards generalisation, it is not claiming that it is AGI, and notes that it said at launch that saturating the benchmark would not represent proof of achieving AGI.

There is a measurement detail underneath that, and it is the part worth carrying away. The 99.9 percent came from a harness supplied by the model’s own developer, which preserves the system’s reasoning state between moves. On the ARC Prize Foundation’s neutral, provider-independent harness, the same model on the same benchmark scored 62.7 percent. Both numbers were published; only one of them travelled.

Two scores 37 points apart, differing only in how the test was administered, is a familiar problem to anyone who works with cognitive assessment. It is why standardised administration is part of the test rather than an afterthought, and it is a large part of why benchmark results move so fast.

What human intelligence testing already learned about this

The problem AGI research is running into now is the one that produced modern psychometrics. A raw count of correct answers tells you almost nothing on its own. To become a score it needs three things, and machine evaluation currently has none of them.

  • A defined population. A score is a position among some group. There is no population of machines and no obvious candidate for one.
  • A norming sample. The test administered under standard conditions to a representative sample, which is what converts a raw count into a scale position.
  • Standardised administration. Fixed conditions, so that two results mean the same thing. The 62.7 versus 99.9 split is exactly what its absence looks like.

Human testing has all three and is still careful: a well-reported score comes with a confidence interval, because measurement error is real and a single number overstates precision. That is the discipline behind reporting a score as a range, and it is conspicuously missing from benchmark figures quoted to one decimal place.

So how close are we?

Honestly: unanswerable as posed, and that is not evasion. Without an agreed definition and an agreed test, the distance to AGI has no units. What can be tracked is narrower and more useful — which specific abilities have fallen, which have not, and how quickly the boundary moves. On that basis the last two years have been remarkable and the remaining gaps are real, particularly around acting in the physical world and learning from very little evidence.

The claim to treat sceptically is not that progress is fast. It is that any single result settles the question. For where machines currently stand against people task by task, the domain-by-domain comparison is the companion to this article; for what happens to your own thinking as you use these systems, the evidence on cognitive debt is more concrete than any forecast. And if the underlying interest is in how general ability gets measured at all, sitting a properly normed reasoning test shows you what a century of solving this problem produced.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

What Is ChatGPT’s IQ Score?

Scores & Scales

ChatGPT IQ Score: Why the Numbers Disagree

Search for ChatGPT’s IQ score and you will find 155, 136, 116 and under 100, all reported seriously, sometimes for the same model. They are not contradictions to be resolved. They are different tests, scored different ways, and the spread between them is the most informative thing about the whole exercise.

Chart of reported AI IQ figures from four different tests, ranging from below 100 on a culture-fair test to 155 on the verbal subtests of the Wechsler scale, showing how far the same systems spread across tests

There is no single ChatGPT IQ score. Published figures for the same family of systems range from slightly below 100 on one culture-fair test to 155 on the verbal half of a clinical instrument, with 116 and 136 reported in between. Every one of those numbers was produced honestly. They disagree because they come from different tests, administered under different conditions, and because none of them is an IQ in the sense the word has when a psychologist uses it.

That spread is worth understanding, because the same forces distort the score a person gets from an online test.

What IQ scores has ChatGPT actually been given?

Four results account for most of what circulates, and they are not measuring comparable things.

  • 155 on the Wechsler verbal subtests. A clinical psychologist administered parts of the Wechsler Adult Intelligence Scale and reported a verbal IQ of 155, above 99.9 percent of the American standardisation sample. Only the verbal subtests could be given: the performance subtests need eyes, ears and hands.
  • Under 100 on a culture-fair test. On the Worldwide IQ Test, a non-verbal assessment designed to minimise language and cultural loading, GPT-4o scored slightly below the population average of 100.
  • 116 and 136 on two other tests. OpenAI’s o3 model was reported at 116 on one assessment and at 136 on the public Mensa Norway test, which would place it above roughly 98 percent of people if it were a person.
  • Up to 151 on tracked weekly testing. The TrackingAI project has been administering the Mensa Norway test to frontier models every week; by September 2026 the top of that chart had reached 151.

A system cannot be simultaneously in the top 0.1 percent and below average. What varies is the test.

Why does the same system score 155 and under 100?

The Wechsler verbal subtests reward stored knowledge: vocabulary, general information, verbal similarities. That is the closest thing to a language model’s home ground, and the result reflects it. The culture-fair test does the opposite. It strips out language and prior knowledge deliberately and asks for pattern completion on abstract figures, which is precisely the ability these systems have found hardest.

In a person, this rarely happens, because human abilities correlate: someone with a top-percentile vocabulary usually does well on matrices too. In a machine there is no such tie between the two, so the choice of test decides the answer. The same principle explains why verbal and non-verbal index scores can diverge sharply in a human profile too, and why a single full-scale number can hide it.

Chart of reported AI IQ figures from four different tests, ranging from below 100 on a culture-fair test to 155 on the verbal subtests of the Wechsler scale, showing how far the same systems spread across tests
Chart of reported AI IQ figures from four different tests, ranging from below 100 on a culture-fair test to 155 on the verbal subtests of the Wechsler scale, showing how far the same systems spread across tests

How do you even give an IQ test to a chatbot?

Less straightforwardly than the headline numbers suggest, and the administration details do a lot of work. The Mensa Norway test behind most published AI figures is a set of 35 visual-pattern puzzles. A model that cannot see is given those puzzles described in words; a model that can see is given the original images. TrackingAI runs both variants weekly and reports the average of the last seven administrations rather than a single sitting.

That produces something genuinely instructive. The same underlying system often appears twice on the chart, once reading a description and once looking at the picture, and the two scores are not the same. One model scored 133 when the puzzles were described to it and 136 when it could see them. Another scored 113 reading and 108 looking. The direction is not even consistent.

  • The format of administration moved the score by several points without anything about the system changing.
  • Averaging seven sittings hides how much any single sitting varies, which for a person is exactly what a confidence interval is for.
  • A test built to be visual becomes a different test when it is read aloud, in the same way a timed test becomes a different test when the clock is removed.

Human testing treats this as fundamental rather than incidental. Standardised administration — same instructions, same time limit, same materials — is part of the instrument, because two scores only mean the same thing if they were produced the same way. Almost none of the AI figures in circulation were.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now!

Secure & encryptedInstant results10–20 minutes

What happens when the test has never been published?

This is the most revealing comparison available, and it comes from the same project that produces the headline numbers. TrackingAI runs two tests on every model it follows. One is Mensa Norway, a public online test that has been on the internet for years and is therefore in the training data. The other is an offline test written by a Mensa member that, in the project’s own words, has never been on the public internet and is in no AI training data.

Reading the underlying chart data on 5 September 2026, 27 models carried a score on both tests. Across those 27, scores on the public test averaged about 13.5 points higher than scores on the unpublished one — close to a full standard deviation on the usual IQ scale. The public test topped out at 151; the unpublished test at 136.

  • One widely used model scored 137 on the public test and 113 on the unpublished one, a 24-point gap.
  • Another scored 145 publicly and 127 privately.
  • A third scored 108 publicly and 88 privately.
  • A small number moved the other way, which is a useful reminder that this is a pattern rather than a law.

An honest caveat belongs here, because it is the kind this site insists on. Two different tests are not directly comparable: they have different norming samples, different difficulty and different item formats, and the unpublished test’s construction has not been released. The gap is consistent with the public test having leaked into the training data, but it does not prove it on its own. What can be said plainly is that models score substantially higher on the test they have almost certainly seen.

The human parallel is exact, and it is the reason serious tests are kept out of circulation. Someone who has worked through a test before scores higher on it the second time without having become any smarter. That is the practice effect, and it is a measured, predictable thing rather than a suspicion.

Is an AI IQ score a real IQ?

No. An IQ is not a mark out of anything. It is a position within a reference population, expressed on a scale with a stated mean — almost always 100 — and a stated standard deviation, usually 15. Producing one requires a norming study in which the test is given to a representative sample of that population under standard conditions.

For a machine, there is no population to be a member of and no norming sample, so the final step simply cannot be taken. What a model produces on an IQ test is a count of correct answers. Calling that count an IQ borrows the authority of a scale it was never placed on. We make the full argument in our piece on why a model scoring 120 is not scoring an IQ.

There is a second problem specific to machines. Standard administration assumes a fixed time, no external help and no second attempt. A model may be run at different effort settings, with or without tools, once or many times, and the reported figure is usually the best of those runs. A human score reported that way would not be accepted either.

What this means for your own score

The lesson transfers directly. A number without a stated scale and a stated reference group is not a result, whoever produced it. When you take a test online, the questions that matter are which population your score is being compared against, what the scale’s standard deviation is, and how much measurement error sits around the figure — a point we cover in the explainer on why a score is a range, not a number.

It also means treating a score you got on a test you had already seen with the same scepticism you would apply to a model’s 151. If you want a number that means something, take a properly normed reasoning test once, cold, and read the result as a range rather than a point. For what the resulting figure does and does not tell you, how to read a test report is the place to start, and the wider human-versus-machine comparison puts the AI numbers in context.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.

Is AI Smarter Than Humans?

Understanding IQ

Is AI Smarter Than Humans? What the 2026 Benchmarks Show

Ask whether AI is smarter than humans and the honest answer depends entirely on the task. Frontier systems now beat expert humans on graduate-level science questions and lose to ordinary people on puzzles built to be unfamiliar. Here is what the 2026 benchmarks actually measure, and why one score never settles it.

Chart comparing machine and human performance across four task types, showing machines ahead on graduate science questions and coding, level on novel puzzle solving, and behind on tasks never seen before

Is AI smarter than humans? On narrow, well-specified tasks the answer in 2026 is often yes, and increasingly by a wide margin. On problems a system has never seen before, described by nobody, with no worked example to copy, people still hold the advantage. The reason both statements are true at once is that "smarter" is not one axis, and the tests that make machines look unbeatable and the tests that stop them cold are measuring genuinely different things.

This article takes the comparison domain by domain, using figures published by the labs and by independent evaluators in 2026, and then explains the structural difference that makes a single answer impossible.

What AI already does better than most people

The clearest machine wins are on tasks with a correct answer, a large body of prior examples, and no requirement to act in the world. On GPQA Diamond, a set of graduate-level biology, chemistry and physics questions written to be hard for people with access to a search engine, OpenAI reported GPT-6 Astra at 96.0 percent in September 2026. Gemini 3.1 Pro sits at 94.3 percent. Both are above the performance of domain experts answering outside their own speciality.

The same pattern holds across coding, long-document retrieval and structured professional work. These are not trick results. They are real, they are reproducible, and they describe abilities that took people years of training to acquire.

  • Recall and synthesis at volume. No person holds the contents of a technical literature in working memory. A model effectively does.
  • Speed. Work that takes an expert a day is returned in minutes, which changes what is worth attempting.
  • Consistency. A model does not get tired on the four-hundredth item, which is exactly where human scorers drift.
  • Breadth of surface knowledge. Competence across far more fields than any individual can maintain.

Where humans still hold the edge

The sharpest counterexample is ARC-AGI-3, a benchmark from the ARC Prize Foundation built specifically to test learning rather than recall. A system is dropped into a small turn-based environment with no instructions and has to work out the goal, the controls and the rules by acting inside it. Crucially, every environment is calibrated on people first: humans solve 100 percent of them, because a task is only admitted once people have shown it can be done.

That calibration is what makes the comparison meaningful, and it is the detail most coverage drops. We set out the full argument in our report on what ARC-AGI-3 measures. The short version: when the novelty is real and the instructions are absent, the gap between an ordinary adult and a frontier system has been enormous.

That gap is now closing fast, and the way it closed is instructive. In September 2026 the ARC Prize Foundation reported GPT-6 Astra at 62.7 percent on its neutral, provider-independent test harness — and 99.9 percent on a harness supplied by the model’s own developer, which preserves the system’s reasoning state between moves. Same model, same benchmark, same week. The scaffolding around the model accounted for most of the difference.

Is AI smarter than humans at learning new things?

This is the question the benchmark was built to ask, and the 2026 answer is genuinely mixed. On action efficiency — how many moves a solver needs to work an unfamiliar environment out — the ARC Prize Foundation found that Astra used fewer actions than the median tested human on 96 percent of the levels it completed, and 51.7 percent fewer actions per level on average. By that measure the machine matched and passed human parity.

Cost tells a different story. The human baseline came from around 500 members of the public, paid roughly 12.78 dollars per game attempted. The model runs that produced those scores cost between 17,332 and 26,098 dollars. The system that learns as efficiently as a person in moves does so at several thousand times the price in resources.

Both numbers are real, and neither alone answers the headline question. That is the pattern to expect from here.

Chart comparing machine and human performance across four task types, showing machines ahead on graduate science questions and coding, level on novel puzzle solving, and behind on tasks never seen before
Chart comparing machine and human performance across four task types, showing machines ahead on graduate science questions and coding, level on novel puzzle solving, and behind on tasks never seen before
Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now!

Secure & encryptedInstant results10–20 minutes

Why smarter is the wrong question

In people, mental abilities correlate. Someone who scores well on vocabulary tends to score well on spatial reasoning and on arithmetic, and the pattern is consistent enough that a single summary number carries real information. That pattern is the reason IQ works at all — we explain the underlying statistics in our explainer on the g factor.

Machines do not show that structure. A system can answer graduate physics correctly and then fail a coloured-grid puzzle that a ten-year-old solves in thirty seconds. Its abilities do not hang together, so no single number summarises them, and any ranking against a person depends entirely on which task you picked. This is not a temporary measurement problem. It is a real difference in how the two kinds of system are built.

  • Human ability is correlated; one score generalises across many tasks.
  • Machine ability is jagged; world-class in one domain, below average in the next, with no reliable pattern.
  • So a comparison needs a task, and any claim without one is not a measurement.

Does a benchmark score mean the same as an IQ score?

No, and the distinction matters more than it sounds. A benchmark result is a raw percentage: items solved out of items attempted. An IQ is not a percentage of anything. It is a position within a reference population, expressed on a scale with a defined mean, almost always 100, and a defined standard deviation, usually 15.

Turning a raw count into that position requires a norming study: the same test, administered under standard conditions, to a representative sample of the population the score will be read against. No such sample exists for machines, and it is not clear what one would even be. So a model can score 96 percent on a science exam without that number converting into any IQ at all. We take that argument apart properly in the article on what ChatGPT’s IQ score really is.

The same logic governs human scores, which is why a raw count on a reasoning test is meaningless until it is placed against a reference sample. That placement is exactly what the IQ percentile calculator does, and it is the step a benchmark percentage has no equivalent for.

When will AI be smarter than humans overall?

Predictions from serious people vary by decades, which is itself the most useful fact about them. Geoffrey Hinton has said he expects machines to surpass human intelligence within about twenty years. Others working on the same systems put it sooner or reject the framing entirely. There is no measurement that would settle the disagreement, because there is no agreed test — the problem we work through in the explainer on what AGI actually means.

What can be said with confidence is narrower. The hardest evaluations still defeat the best systems: on Humanity’s Last Exam, a set of expert-written questions across dozens of fields, the strongest reported result in September 2026 was 65.0 percent, meaning better than a third of the questions remained unanswered. Benchmarks that were supposed to hold for years keep falling, and new ones keep being built because the old ones stop separating anything.

What this means for measuring your own intelligence

None of this changes what a cognitive test does for a person. The reason matrix puzzles sit near the centre of most non-verbal reasoning tests is that they lean as little as possible on what you happen to know and as much as possible on working out a rule from the evidence in front of you. That the same format is what machines found hardest is a point in favour of the format, not against it.

If the comparison has made you curious about your own reasoning rather than a model’s, our IQ test is built around exactly this kind of rule-finding, and the result is reported the way a score has to be reported to mean anything: as a position in a reference population, with a range around it. For what those numbers do and do not predict, the evidence on outcomes is a better guide than any headline about machines.

The durable conclusion is unglamorous. Machines are now better than most people at a growing list of specific things, worse at a shrinking list, and not comparable at all on the single scale the question implies. Anyone offering you one number for it is selling something.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.