IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment
AI & Machine IntelligenceEmerging evidence

What the Artificial Analysis Intelligence Index Actually Measures

Headlines call it an AI intelligence score. It is a weighted average of ten separate evaluations, and the hardest one — an exam built to resist quick solving — is already close to 60% solved.

What the Artificial Analysis Intelligence Index Actually Measures
Illustration generated for IQ Metrics. No photograph is used. IQ Metrics

What you need to know

  • The Artificial Analysis Intelligence Index is a weighted composite of ten separate evaluations grouped into four categories — agentic tasks (30% of the total score), general knowledge (30%), coding (20%) and scientific reasoning (20%) — not a single test, and its methodology has already been revised multiple times, most recently to version 4.3.
  • Humanity’s Last Exam, the index’s hardest single ingredient, is roughly 2,500 graduate-level questions built specifically to resist the quick saturation that retired earlier benchmarks. It is still only one of ten evaluations feeding the composite.
  • Scores on that exam moved from about 2.7% (GPT-4o) to 8% (OpenAI’s o1) to roughly 59% (Claude Fable 5.1, measured in early September 2026) in under two years — a pace that says as much about how fast a benchmark gets partly solved as it does about any one model.
  • Two models can post the same overall index score while differing by dozens of points on any single evaluation underneath it — the same structural fact this site’s own explainer on full-scale IQ versus index scores makes about a human test report.

Every few weeks brings a new headline declaring one company’s AI model "the most intelligent." The evidence behind it is almost always the same kind of thing: a single number from a leaderboard called an intelligence index. Our earlier piece on what one such claim actually checked out to covered a specific announcement. This one opens the number itself — what it is built from, and what gets lost when ten separate evaluations get folded into one score.

Ten evaluations, four categories, one weighted average

The Artificial Analysis Intelligence Index, one of the most widely cited composite scores in AI reporting, combines ten separate evaluations into four weighted categories: agentic tasks at 30% of the total, general knowledge at 30%, coding at 20% and scientific reasoning at 20%, according to the index’s own published methodology as of September 2026. Inside those categories, individual evaluations carry their own sub-weights again — one single agentic evaluation accounts for 15% of the entire index on its own, a bigger share than either of the two evaluations that make up the scientific-reasoning category combined.

  • Agents, 30% of total: AA-Briefcase, GDPval-AA v2, AutomationBench-AA
  • Coding, 20%: Terminal-Bench v4.0, SciCode
  • General, 30%: AA-Omniscience, GDP.pdf, AA-LCR v1.1
  • Scientific Reasoning, 20%: Humanity’s Last Exam, CritPt

That weighting is a choice, not a law of nature. An index built to reward finishing multi-step agentic tasks will rank models differently than one built to reward graduate-level scientific recall, and the current version weights the former more than twice as heavily as the latter. The methodology has already been revised more than once — the version in use in September 2026 is 4.3 — so a model’s rank today is not directly comparable to a rank read off an older version of the same-named index.

Why "agentic" tasks got the biggest single weight

GDPval-AA v2, one of the three agentic evaluations, illustrates what that category actually tests. The underlying GDPval benchmark, released by OpenAI, is built from tasks drawn from the real work of industry professionals — an average of 14 years of experience each — across 44 occupations spanning the top nine sectors that contribute most to U.S. GDP: finance, healthcare, law, software, and others. Its 220-task gold subset asks a model to produce an actual work product — a slide deck, a legal memo, a spreadsheet model, a diagram — graded by having human experts compare the model’s output against a professional’s, rather than checking a single right answer. OpenAI’s own reporting on the benchmark states that frontier models complete these tasks roughly 100 times faster and cheaper than the human professionals whose work anchors the comparison, though speed and cost are not the same thing as quality of output.

That is a meaningfully different thing to measure than a written exam. A model can be excellent at recalling and applying specialist knowledge under exam conditions and comparatively weak at the sustained, multi-step judgment a real deliverable requires, or the reverse. Weighting agentic, real-work evaluations at 30% — the joint-largest share, tied with general knowledge — is Artificial Analysis’s stated bet that economically useful task completion is the more important thing to track. That is a defensible editorial choice. It is also a choice, and a different one would move the leaderboard without any model actually changing.

The hardest ingredient: an exam built to resist being solved

Humanity’s Last Exam is the scientific-reasoning half of the 20% category with the smallest weight, and it exists because the benchmarks before it stopped working. The Center for AI Safety and Scale AI released it publicly in January 2025, built from roughly 2,500 questions spanning more than 100 subjects — about 41% mathematics, with roughly 14% of the set incorporating multi-modal elements such as images or diagrams, and the remainder split across physics, biology, medicine, computer science and the humanities — contributed by close to 1,000 subject-matter experts across more than 500 institutions in 50 countries. Submitted questions were tested against leading AI models first; only ones the models answered incorrectly, or did no better than random guessing on, went through two further rounds of human expert review before making the final set, with a $500,000 prize pool distributed to the strongest submissions. Stanford HAI’s AI Index 2025 report cites the project as a direct response to benchmark saturation — frontier models were already scoring close to the ceiling on the popular benchmarks that came before it.

The name is deliberately dramatic. A model that scores well on it has not matched human ability across the board — the questions are narrow, graduate-level and hand-picked specifically for being hard to answer without deep specialist knowledge, which is a different thing from general competence.
  • GPT-4o (2024): 2.7%
  • Claude 3.5 Sonnet (2024): 4.1%
  • OpenAI o1: 8%
  • Gemini 3.1 Pro and Claude Opus 4.6 (later rounds): roughly 40–50%
  • Claude Fable 5.1 and GPT-6 Astra, text-only subset, measured 3 September 2026: 59.1% and 54.7%

An exam built specifically to be too hard for the models that existed when it launched is now, on its text-only questions, better than half-solved by the model that also tops the composite index it feeds into.

What one composite number hides

A single index score invites a mistake this site has already covered in a different context. Our explainer on full-scale IQ versus index scores makes the point that a single full-scale number is really an average across several separate measurements, and that two of those measurements can pull apart on the same person’s report even when the top-line score looks unremarkable. The same structure sits inside an AI intelligence index: two models can land on an identical overall score while one is well ahead on agentic task completion and meaningfully behind on scientific reasoning, or the reverse. The composite does not show which — and unlike a WISC-V or WAIS report, there is no equivalent of the confidence-interval band this site’s piece on IQ score margin of error covers, printed next to the headline AI figure for a general reader to see.

It is also worth being precise about what these scores are not. Our piece on why an AI’s score on an IQ-style question set is not an IQ covers the deeper reason: an IQ score is anchored to a norm — the distribution of scores in an actual human reference population, expressed as a percentile against real people who sat the same test under the same conditions. An intelligence index score has no such population behind it. It is a percentage or a points total on a fixed set of questions, comparable only to another model’s score on that identical set, not to a percentile ranking against anyone.

Why the same two models get reported differently by different sources

None of this means the underlying numbers are arbitrary. Artificial Analysis publishes its methodology and reports a 95% confidence interval under plus-or-minus one point from repeated testing, tighter than most of the gaps argued over in headlines. What it means is that "intelligence index" is a brand name for one organization’s particular weighting choice, and a different aggregator weighting the same ten-odd evaluations differently would produce a different leaderboard from identical underlying scores. Our piece on why the same model gets different benchmark scores from different labs covers a related but separate cause of disagreement — grading-rule choices inside a single benchmark, which alone can move one model’s reported score by dozens of points.

Your own number

Where would your own score land?

Curious what an actual percentile-anchored score looks like? Take the full IQ assessment or run a number through our percentile calculator to see how a real reference population turns a raw score into a rank.

Find your IQ score now!
Secure & encryptedInstant results10–20 minutes

Next time a leaderboard number gets reported as "the most intelligent AI," the useful question is not which company published it. It is which ten evaluations are sitting underneath the average, and how much weight each one was given before anyone saw the final digit. Our own IQ test reports a percentile against a real population for exactly this reason — a single number is only as meaningful as what stands behind it.

Common questions

What is the Artificial Analysis Intelligence Index?

A composite score published by Artificial Analysis that combines ten separate AI evaluations — covering agentic tasks, general knowledge, coding and scientific reasoning — into one weighted average, used to compare language models such as GPT, Claude and Gemini.

What is Humanity’s Last Exam?

A benchmark of roughly 2,500 graduate-level questions across more than 100 subjects, built by the Center for AI Safety and Scale AI and released in January 2025 specifically to resist the quick saturation that had made earlier AI benchmarks less useful.

Does a high score on an AI intelligence index mean the model has a high IQ?

No. An IQ score is anchored to a human reference population and expressed as a percentile against people who took the same test. An AI intelligence index score is a percentage or points total on a fixed question set, comparable only to another model’s score on that same set.

Why do AI model rankings change so often?

Composite indices get their methodology revised — evaluations added, weights changed — and the underlying evaluations keep shifting as newer models catch up to older ones, so a leaderboard position under one version of an index is not directly comparable to a later version.

Sources for this story

  1. Intelligence Index methodology, version 4.3 — Artificial Analysis
  2. Humanity’s Last Exam benchmark release — Center for AI Safety and Scale AI (2025)
  3. AI Index 2025 Annual Report, on benchmark saturation — Stanford Institute for Human-Centered AI (HAI)
  4. GDPval: evaluating AI model performance on real-world economically valuable tasks — OpenAI (2025)
  5. Humanity’s Last Exam leaderboard and text-only subset scores — Epoch AI

Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.

Share this story

Know someone who keeps seeing this number quoted without the scale it was measured on? Send it to them — it takes one tap.

Filed under#ai benchmarks#scores and scales#confidence intervals#study quality

Read the research.
Then find your own number.

Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.

Start IQ Test
Secure & encryptedInstant results10–20 minutes