What Is AGI, and How Would Anyone Measure It?
Artificial general intelligence means a system that can learn and apply knowledge across the full range of tasks a person can, rather than excelling at a narrow set. The trouble is that the major labs each publish a different definition, none of them converts into a test, and so claims that AGI has arrived cannot be checked either way.

AGI stands for artificial general intelligence: an AI system with the ability to learn, reason and apply knowledge across the full range of tasks a person can handle, rather than performing brilliantly in one narrow domain. That much is broadly agreed. Almost nothing after it is.
The organisations building these systems publish materially different definitions, none of those definitions specifies a test, and there is no reference population against which a result could be scored. So when a launch announcement is described as the arrival of AGI, there is no procedure available to check the claim. That is a measurement problem, and human intelligence testing solved a version of it a century ago.
Why nobody agrees on what AGI means
The published definitions differ on what would count, and the differences are not cosmetic.
- OpenAI: autonomous systems that outperform humans at most economically valuable work.
- Google: AI that is at least as capable as humans at most cognitive tasks.
- Microsoft: the point at which an AI can match human performance at all tasks.
- ARC Prize Foundation: a system’s ability to acquire any skill a human can, as efficiently as a human can.
Read them side by side and the gaps open up. OpenAI’s is economic and could in principle be satisfied without the system ever learning anything new. Microsoft’s says all tasks, which is a far higher bar than most. The ARC Prize definition is the only one that puts learning efficiency at the centre, which is also the only version that resembles what psychologists mean by general ability.
The disagreement is not merely academic, because each definition implies a different finish line and a different set of evidence. An economic definition is settled by labour-market data. A cognitive-task definition is settled by evaluations. A learning-efficiency definition is settled by putting a system in front of something genuinely unfamiliar and counting how much evidence it needs. A system could satisfy one and fail another on the same day, which is roughly where things stand.
A term that means four things does not support a yes-or-no answer. It supports four of them.
What is the difference between AI, AGI and superintelligence?
The three terms describe a ladder, and most confusion comes from using them interchangeably.
- Narrow AI is everything in use today: systems that perform specific tasks, sometimes far better than people, without the ability to transfer that competence to an unrelated problem. A model that writes excellent code and cannot fold a towel is narrow, however impressive the code.
- AGI would match human ability across the full breadth of tasks rather than in a slice of them. Breadth is the operative word: the claim is about range, not peak performance.
- Superintelligence would exceed human ability across that same full range. It is a claim about what comes after AGI, and is entirely hypothetical.
The distinction that matters for measurement is the first one. Narrow ability is straightforward to test, because you can specify the task. Breadth is not, because you would have to specify every task — which is the problem the next section is about.
How would anyone actually measure AGI?
The most serious attempt starts from a 2019 paper by François Chollet called "On the Measure of Intelligence", which argued that nearly every AI benchmark of the period measured skill at a task — and that skill at a task can simply be bought with enough data and training. What a benchmark should measure instead, on that argument, is how efficiently a system picks up a skill it was never trained on.
That produced the ARC-AGI series, built around puzzles with no vocabulary, no cultural reference and no arithmetic beyond counting, where every task has a different rule and nothing carries over. The distinction it targets is the one human testing already draws between fluid reasoning — working out something new — and crystallized ability, which is what you have accumulated. The general factor in human ability leans most heavily on the first.
Older proposals still get cited and are worth knowing. The Turing test asked whether a machine could be told apart from a person in conversation. Marvin Minsky suggested in 1970 that a machine should be able to read Shakespeare, grease a car, play office politics and tell a joke. The coffee test, associated with Steve Wozniak, proposes that AGI arrives when a machine can walk into an unfamiliar house and make a pot of coffee. Each captures something. None produces a score.

Where would your own score land?
Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.
Find your IQ score now! →Did GPT-6 Astra just reach AGI?
In September 2026 OpenAI announced GPT-6 Astra, and reported 99.9 percent on ARC-AGI-3 — the interactive benchmark built specifically to resist this kind of result. Coverage framed it as the arrival of the AGI era.
The organisation that owns the benchmark disagrees, in writing. The ARC Prize Foundation’s own analysis states plainly that while it believes the system represents meaningful progress towards generalisation, it is not claiming that it is AGI, and notes that it said at launch that saturating the benchmark would not represent proof of achieving AGI.
There is a measurement detail underneath that, and it is the part worth carrying away. The 99.9 percent came from a harness supplied by the model’s own developer, which preserves the system’s reasoning state between moves. On the ARC Prize Foundation’s neutral, provider-independent harness, the same model on the same benchmark scored 62.7 percent. Both numbers were published; only one of them travelled.
Two scores 37 points apart, differing only in how the test was administered, is a familiar problem to anyone who works with cognitive assessment. It is why standardised administration is part of the test rather than an afterthought, and it is a large part of why benchmark results move so fast.
What human intelligence testing already learned about this
The problem AGI research is running into now is the one that produced modern psychometrics. A raw count of correct answers tells you almost nothing on its own. To become a score it needs three things, and machine evaluation currently has none of them.
- A defined population. A score is a position among some group. There is no population of machines and no obvious candidate for one.
- A norming sample. The test administered under standard conditions to a representative sample, which is what converts a raw count into a scale position.
- Standardised administration. Fixed conditions, so that two results mean the same thing. The 62.7 versus 99.9 split is exactly what its absence looks like.
Human testing has all three and is still careful: a well-reported score comes with a confidence interval, because measurement error is real and a single number overstates precision. That is the discipline behind reporting a score as a range, and it is conspicuously missing from benchmark figures quoted to one decimal place.
So how close are we?
Honestly: unanswerable as posed, and that is not evasion. Without an agreed definition and an agreed test, the distance to AGI has no units. What can be tracked is narrower and more useful — which specific abilities have fallen, which have not, and how quickly the boundary moves. On that basis the last two years have been remarkable and the remaining gaps are real, particularly around acting in the physical world and learning from very little evidence.
The claim to treat sceptically is not that progress is fast. It is that any single result settles the question. For where machines currently stand against people task by task, the domain-by-domain comparison is the companion to this article; for what happens to your own thinking as you use these systems, the evidence on cognitive debt is more concrete than any forecast. And if the underlying interest is in how general ability gets measured at all, sitting a properly normed reasoning test shows you what a century of solving this problem produced.
Keep reading
All articles →
Understanding IQIs AI Smarter Than Humans?
Ask whether AI is smarter than humans and the honest answer depends entirely on the task. Frontier systems now beat expert humans on graduate-level science questions and lose to ordinary people on puzzles built to be unfamiliar. Here is what the 2026 benchmarks actually measure, and why one score never settles it.
Mind & Everyday LifeDoes Using AI Make You Dumber?
The claim that AI is making us stupider has real research behind it and is still routinely overstated. No study has measured a drop in intelligence. What has been measured is lower mental engagement during a delegated task and weaker memory for work you did not do yourself, which is a narrower finding and a more useful one.
Research & EvidenceGardner's Multiple Intelligences
Gardner's multiple intelligences theory is one of the most influential ideas in education and almost entirely absent from the tests psychologists use. That gap is not stubbornness. It comes down to one stubborn statistical finding, one missing instrument, and a real insight the theory carries.
