How machine reasoning is measured, where it beats people, where it still fails, and what any of it says about human intelligence.
The sturdiest results are narrow and decades old. The 2025 studies that drew the headlines rest on self-reports and small preprints — and none of them measured intelligence at all.
IQ scales are defined by a human reference sample, and a language model is not in it. Placing a machine on that scale is not a hard measurement problem — it is a category error.
A bad night does not lower intelligence. It reliably lowers the things a timed test happens to measure most closely — sustained attention, working memory and speed — while leaving vocabulary and general knowledge comparatively intact.
Two people can perform identically on two different online tests and walk away with scores twenty points apart. The scale, the comparison group and the margin of error are what separate an assessment from a quiz.
Average scores climbed for decades, then flattened and in places slipped. None of it shows on a score report, because every test is reset so the average is 100 again — which is why a 1990 score is not a 2026 score.
An IQ figure is meaningless without the scale it was measured on and the range around it. Every story here reports all three, or says plainly that the source did not.
Each story ends with the studies, datasets and documents it draws on, named and attributed, so you can go and read them yourself.
Spotted an error? Write to corrections@iqmetrics.org. Corrections are made on the story and noted at the bottom of it — never quietly.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →