IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment
AI & Machine IntelligenceEmerging evidence

Plausible Nonsense: Only 4 of 60 AI Models Cleared a Reasoning Benchmark

A PNAS paper published on 9 September 2026 tested 60 large language models against human deliberation. Only four consistently cleared its benchmark, while the rest could still sound reasonable. Here is what the abstract does and does not establish.

Plausible Nonsense: Only 4 of 60 AI Models Cleared a Reasoning Benchmark
Illustration generated for IQ Metrics. No photograph is used. IQ Metrics

What you need to know

  • Veri and Kreia Umbelino (PNAS 123(37), e2600126123, published 9 September 2026) compared 60 large language models with human deliberation across nine policy scenarios using the Deliberative Reason Index. Only four models consistently exceeded the study’s permutation-based null benchmark.
  • The paper is not an IQ test and does not rate models on intelligence. It asks whether the reasons a model gives line up with its preferences the way they do among people who deliberated together.
  • Fluency is not the signal. The authors say outputs can appear coherent and persuasive even when that alignment is absent, and urge assessing a model’s deliberative reasoning before it is used in governance.
  • An earlier benchmark by the same authors, described in a 2026 arXiv paper, points the same way: across 54 models and 526 human responses, the average index was 0.18 for models against 0.34 for people.

A paper in the Proceedings of the National Academy of Sciences (PNAS) on 9 September 2026 asks a narrower question than whether AI is smart. It asks whether the reasons a language model gives line up with the choices it makes in the way they do among people who have deliberated together. Francesco Veri and Gustavo Kreia Umbelino tested 60 large language models (LLMs) against human deliberation across nine policy scenarios, using a measure called the Deliberative Reason Index (DRI). Only four models consistently exceeded the study’s permutation-based null benchmark. The rest did not, even though their outputs can still read as coherent and persuasive, which is the paper’s point: fluency and reasoning structure are different things.

We read the paper’s abstract and its publication record. The full text sits behind PNAS’s site and could not be opened for this article, so we do not name the four models or quote per-model scores.

What did the PNAS study actually test?

The setting is democratic governance. The authors note that LLMs are entering democratic contexts as instruments of governance, where problems are ill-structured: marked by ambiguity and contestation rather than a single correct answer. Such problems, they argue, demand more than factual precision. They call for intersubjective reasoning, meaning context-sensitive judgments that other people can understand and publicly accept. So the test is not a quiz with a key. Sixty models were compared with human deliberation across nine policy scenarios, using the DRI.

What is the Deliberative Reason Index?

The DRI measures intersubjective consistency. A later arXiv paper by Maurice Flechtner (2608.10186, 10 August 2026) describes it in terms of two correlations for pairs of participants, one across the considerations they weigh and one across the preferences they hold, and defines consistency as the similarity of the two. High consistency means that when two participants share or diverge in their reasoning about considerations, this is reflected proportionally in their shared or divergent preferences. In plain words, people who reason alike should end up choosing alike, and people who reason differently should choose differently. A reasoner whose reasons and choices do not track each other that way is not deliberating coherently, however good each sentence sounds.

What does “plausible nonsense” mean here?

The phrase names a gap the abstract states directly: outputs can appear reasonable even when the alignment with human reason-giving is absent, and the gap between surface plausibility and deliberative coherence is the reason for caution. That answers a question the title invites. The paper does not test whether a model can spot nonsense in other people’s text. It tests whether the model’s own pattern of reasons and preferences behaves like a human deliberator’s. Because fluent prose is what these systems are built to produce, fluency cannot be the test.

The paper’s target is the gap between sounding reasonable and reasoning coherently.

What did the same authors find before?

This is not the first benchmark from the team. A separate analysis by Kreia Umbelino and Veri (2025), as described in Flechtner’s paper, evaluated 54 off-the-shelf LLMs against 526 post-deliberation human DRI responses across 24 cases spanning 19 topics. Humans scored higher on average, 0.34 against 0.18 for the models (p < 0.0001), and outperformed the models in 19 of the 24 cases. At model level, humans consistently outperformed 23 of the 54 models, while the remaining majority performed on par with humans on average. The models were tested as released, without fine-tuning or access to deliberation transcripts, so Flechtner treats the result as a baseline rather than a ceiling. The PNAS paper uses 60 models and nine scenarios, so the two sets of numbers are different analyses and should not be merged.

How is this different from an IQ-style benchmark?

An IQ test scores right and wrong answers against a norm group and reports where the taker falls, in percentile terms. This benchmark scores something else: whether reasons and preferences cohere across many decisions. A model can do well on answer-based tests and still fall short on structure, and the reverse. That is one reason we argue that an AI’s score on an IQ-style test is not an IQ, and why AI benchmark scores disagree when the same models are measured in different ways.

Four questions to ask of any AI reasoning headline

  • What is the baseline? A chance-level benchmark, a human benchmark and a previous model are three different bars.
  • What is being scored? Answers, or the structure behind the answers.
  • How many settings? Consistent results across nine scenarios are a stronger claim than one result.
  • Which versions and dates? Models change every few months, so a result belongs to a snapshot.
In general, a permutation-based null benchmark is built by shuffling the data many times to see what chance alone produces. The abstract does not spell out how this study built its benchmark, so clearing it should be read as doing better than the study’s chance baseline, not as equalling a person.

What does this mean for reading AI “IQ” claims?

It is a reminder that one number rarely captures reasoning. Human IQ is a percentile on a norm group, with a stated mean, standard deviation and margin of error; the percentile calculator shows how that works. AI results are often reported as a single leaderboard score without any of that, as in our look at what the intelligence index says. If you want a human reference point of your own, take the IQ test and read the result as a range.

Your own number

Where would your own score land?

Want a human benchmark of your own? Take the IQ Metrics test and read your score as a percentile.

Find your IQ score now! →
Secure & encryptedInstant results10–20 minutes

The takeaway is narrow and useful. In this study, most of 60 models did not consistently clear a bar for human-like reason-giving, and the authors want that checked before models are deployed in governance. It says little about models that were not tested and nothing about a model’s general intelligence.

Common questions

What did the 2026 PNAS study find about LLMs and reasoning?

Veri and Kreia Umbelino tested 60 large language models against human deliberation across nine policy scenarios using the Deliberative Reason Index. Only four models consistently exceeded the study’s permutation-based null benchmark for alignment with human patterns of reason-giving.

What is the Deliberative Reason Index?

It measures intersubjective consistency: whether people who share or diverge in their reasoning about considerations also share or diverge, proportionally, in their preferences. It compares two correlations across pairs of participants.

Can AI spot plausible nonsense?

That is not what this paper tests. It measures whether a model’s own reasons and preferences hang together like a human deliberator’s. Its authors say outputs can look coherent and persuasive even when that alignment is absent.

Is this an IQ test for AI?

No. It does not report an IQ or a percentile on a human norm group. It scores the consistency of reasons and preferences against human deliberation, so it should not be read as a general intelligence rating.

Which four models passed the benchmark?

The abstract does not name them, and the full text could not be opened for this article, so we do not name them either.

Sources for this story

  1. Plausible nonsense and deliberative reasoning: Benchmarking LLMs against human judgment (F. Veri and G. Kreia Umbelino), doi 10.1073/pnas.2600126123 — Proceedings of the National Academy of Sciences 123(37), e2600126123, published online 9 September 2026 (abstract read; full text not accessed)
  2. Crossref metadata record for doi 10.1073/pnas.2600126123 — Crossref
  3. The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse (M. Flechtner), arXiv 2608.10186v1 — arXiv, 10 August 2026
  4. Kreia Umbelino and Veri (2025), benchmark of 54 off-the-shelf LLMs, as summarised in the arXiv paper above (original not opened) — Cited in Flechtner, arXiv 2608.10186

Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.

Share this story

Know someone who keeps seeing this number quoted without the scale it was measured on? Send it to them — it takes one tap.

Filed under#ai benchmarks#fluid reasoning#study quality#scores and scales

Read the research.
Then find your own number.

Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.

Start IQ Test →
Secure & encryptedInstant results10–20 minutes