Headlines call it an AI intelligence score. It is a weighted average of ten separate evaluations, and the hardest one — an exam built to resist quick solving — is already close to 60% solved.

Every few weeks brings a new headline declaring one company’s AI model "the most intelligent." The evidence behind it is almost always the same kind of thing: a single number from a leaderboard called an intelligence index. Our earlier piece on what one such claim actually checked out to covered a specific announcement. This one opens the number itself — what it is built from, and what gets lost when ten separate evaluations get folded into one score.
The Artificial Analysis Intelligence Index, one of the most widely cited composite scores in AI reporting, combines ten separate evaluations into four weighted categories: agentic tasks at 30% of the total, general knowledge at 30%, coding at 20% and scientific reasoning at 20%, according to the index’s own published methodology as of September 2026. Inside those categories, individual evaluations carry their own sub-weights again — one single agentic evaluation accounts for 15% of the entire index on its own, a bigger share than either of the two evaluations that make up the scientific-reasoning category combined.
That weighting is a choice, not a law of nature. An index built to reward finishing multi-step agentic tasks will rank models differently than one built to reward graduate-level scientific recall, and the current version weights the former more than twice as heavily as the latter. The methodology has already been revised more than once — the version in use in September 2026 is 4.3 — so a model’s rank today is not directly comparable to a rank read off an older version of the same-named index.
GDPval-AA v2, one of the three agentic evaluations, illustrates what that category actually tests. The underlying GDPval benchmark, released by OpenAI, is built from tasks drawn from the real work of industry professionals — an average of 14 years of experience each — across 44 occupations spanning the top nine sectors that contribute most to U.S. GDP: finance, healthcare, law, software, and others. Its 220-task gold subset asks a model to produce an actual work product — a slide deck, a legal memo, a spreadsheet model, a diagram — graded by having human experts compare the model’s output against a professional’s, rather than checking a single right answer. OpenAI’s own reporting on the benchmark states that frontier models complete these tasks roughly 100 times faster and cheaper than the human professionals whose work anchors the comparison, though speed and cost are not the same thing as quality of output.
That is a meaningfully different thing to measure than a written exam. A model can be excellent at recalling and applying specialist knowledge under exam conditions and comparatively weak at the sustained, multi-step judgment a real deliverable requires, or the reverse. Weighting agentic, real-work evaluations at 30% — the joint-largest share, tied with general knowledge — is Artificial Analysis’s stated bet that economically useful task completion is the more important thing to track. That is a defensible editorial choice. It is also a choice, and a different one would move the leaderboard without any model actually changing.
Humanity’s Last Exam is the scientific-reasoning half of the 20% category with the smallest weight, and it exists because the benchmarks before it stopped working. The Center for AI Safety and Scale AI released it publicly in January 2025, built from roughly 2,500 questions spanning more than 100 subjects — about 41% mathematics, with roughly 14% of the set incorporating multi-modal elements such as images or diagrams, and the remainder split across physics, biology, medicine, computer science and the humanities — contributed by close to 1,000 subject-matter experts across more than 500 institutions in 50 countries. Submitted questions were tested against leading AI models first; only ones the models answered incorrectly, or did no better than random guessing on, went through two further rounds of human expert review before making the final set, with a $500,000 prize pool distributed to the strongest submissions. Stanford HAI’s AI Index 2025 report cites the project as a direct response to benchmark saturation — frontier models were already scoring close to the ceiling on the popular benchmarks that came before it.
An exam built specifically to be too hard for the models that existed when it launched is now, on its text-only questions, better than half-solved by the model that also tops the composite index it feeds into.
A single index score invites a mistake this site has already covered in a different context. Our explainer on full-scale IQ versus index scores makes the point that a single full-scale number is really an average across several separate measurements, and that two of those measurements can pull apart on the same person’s report even when the top-line score looks unremarkable. The same structure sits inside an AI intelligence index: two models can land on an identical overall score while one is well ahead on agentic task completion and meaningfully behind on scientific reasoning, or the reverse. The composite does not show which — and unlike a WISC-V or WAIS report, there is no equivalent of the confidence-interval band this site’s piece on IQ score margin of error covers, printed next to the headline AI figure for a general reader to see.
It is also worth being precise about what these scores are not. Our piece on why an AI’s score on an IQ-style question set is not an IQ covers the deeper reason: an IQ score is anchored to a norm — the distribution of scores in an actual human reference population, expressed as a percentile against real people who sat the same test under the same conditions. An intelligence index score has no such population behind it. It is a percentage or a points total on a fixed set of questions, comparable only to another model’s score on that identical set, not to a percentile ranking against anyone.
None of this means the underlying numbers are arbitrary. Artificial Analysis publishes its methodology and reports a 95% confidence interval under plus-or-minus one point from repeated testing, tighter than most of the gaps argued over in headlines. What it means is that "intelligence index" is a brand name for one organization’s particular weighting choice, and a different aggregator weighting the same ten-odd evaluations differently would produce a different leaderboard from identical underlying scores. Our piece on why the same model gets different benchmark scores from different labs covers a related but separate cause of disagreement — grading-rule choices inside a single benchmark, which alone can move one model’s reported score by dozens of points.
Curious what an actual percentile-anchored score looks like? Take the full IQ assessment or run a number through our percentile calculator to see how a real reference population turns a raw score into a rank.
Find your IQ score now! →Next time a leaderboard number gets reported as "the most intelligent AI," the useful question is not which company published it. It is which ten evaluations are sitting underneath the average, and how much weight each one was given before anyone saw the final digit. Our own IQ test reports a percentile against a real population for exactly this reason — a single number is only as meaningful as what stands behind it.
A composite score published by Artificial Analysis that combines ten separate AI evaluations — covering agentic tasks, general knowledge, coding and scientific reasoning — into one weighted average, used to compare language models such as GPT, Claude and Gemini.
A benchmark of roughly 2,500 graduate-level questions across more than 100 subjects, built by the Center for AI Safety and Scale AI and released in January 2025 specifically to resist the quick saturation that had made earlier AI benchmarks less useful.
No. An IQ score is anchored to a human reference population and expressed as a percentile against people who took the same test. An AI intelligence index score is a percentage or points total on a fixed question set, comparable only to another model’s score on that same set.
Composite indices get their methodology revised — evaluations added, weights changed — and the underlying evaluations keep shifting as newer models catch up to older ones, so a leaderboard position under one version of an index is not directly comparable to a later version.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
OpenAI launched GPT-6 Astra on 3 September with that phrase. The one composite intelligence measure in the announcement's own comparison table puts it in fourth place.
OpenAI and Anthropic published figures for the same rival models within four days of each other. They do not match, and the reasons are the ones psychometrics named a century ago.

Language models that do well on one benchmark tend to do well on the others, and published psychometric analyses have extracted a dominant common factor from those scores. Whether that factor is anything like human g depends on a question about what is varying between the models.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →