How machine reasoning is measured, where it beats people, where it still fails, and what any of it says about human intelligence.

Headlines call it an AI intelligence score. It is a weighted average of ten separate evaluations, and the hardest one — an exam built to resist quick solving — is already close to 60% solved.

Language models that do well on one benchmark tend to do well on the others, and published psychometric analyses have extracted a dominant common factor from those scores. Whether that factor is anything like human g depends on a question about what is varying between the models.
GPT-6 Astra scored 62.7 per cent and 99.9 per cent on the same benchmark in the same week. The difference is not the model. It is how the test was administered.
Machines have been writing reasoning items since long before the current wave of language models, and the research is clear about where the difficulty lies. Generating a plausible question is the easy half. Knowing how hard it is remains the expensive half.
Frontier systems now clear reasoning benchmarks that defeated them a year ago. Put them in front of puzzles screened so that ordinary people solve every one, and the best models finish under one per cent.
The sturdiest results are narrow and decades old. The 2025 studies that drew the headlines rest on self-reports and small preprints — and none of them measured intelligence at all.
IQ scales are defined by a human reference sample, and a language model is not in it. Placing a machine on that scale is not a hard measurement problem — it is a category error.
An IQ figure is meaningless without the scale it was measured on and the range around it. Every story here reports all three, or says plainly that the source did not.
Each story ends with the studies, datasets and documents it draws on, named and attributed, so you can go and read them yourself.
Spotted an error? Write to corrections@iqmetrics.org. Corrections are made on the story and noted at the bottom of it — never quietly.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →