IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

Has AI Passed the Turing Test? What the 2026 PNAS Study Found

A peer-reviewed PNAS paper says GPT-4.5, given a humanlike persona, was picked as the human in 73% of five-minute chats. Without the persona it managed 36%, and the authors say the test measures humanlikeness, not intelligence.

Has AI Passed the Turing Test? What the 2026 PNAS Study Found
Illustration generated for IQ Metrics. No photograph is used. IQ Metrics

What you need to know

  • Jones and Bergen's three-party Turing test, published in PNAS on 19 May 2026, found GPT-4.5 with a persona prompt judged human in 73% of 5-minute games, LLaMa-3.1-405B in 56%, and ELIZA, a 1960s rules-based chatbot, in 23%. Our 95% interval for the 73% is about 67% to 78%.
  • The result depends on the prompt. The same GPT-4.5 without the persona scored 36%, and neither it nor GPT-4o did significantly better than ELIZA. A 15-minute replication with GPT-5 gave 59%, borderline above chance after correction, and all three studies come from one research group.
  • Interrogators mostly judged style, not smarts: in 12% of games they quizzed knowledge or reasoning, and one of the most accurate cues was a witness not knowing something an AI should know. Turing's own 1950 figure was a prediction, not a pass mark.
  • A Turing result is a win rate, not an IQ score. We found no study that compares Turing-test judgments with IQ-style test results, and the PNAS authors say the test is a measure of humanlikeness.

Yes, in one specific sense. A peer-reviewed study published in PNAS in May 2026 reports that GPT-4.5, told to act like a young, introverted person who uses slang, was judged to be the human in 73% of five-minute, three-way conversations, more often than the real human it was paired with. The same model without that persona prompt scored 36%, and a 15-minute follow-up with GPT-5 gave 59%. A pass here means a win rate not significantly below 50%. It is not an IQ score and not a test score at all. This guide sets out what the study did, what the numbers do and do not show, what interrogators actually asked, and why the authors say the test measures humanlikeness, not intelligence.

What did the 2026 PNAS Turing test study find?

Cameron Jones (Stony Brook University) and Benjamin Bergen (UC San Diego) ran a standard three-party Turing test: an interrogator chats in text, at the same time, with one human and one machine, then says which is the human. The paper, Large language models pass a standard three-party Turing test, appeared in PNAS volume 123, issue 21, online on 19 May 2026. It began as a March 2025 preprint and was accepted on 27 March 2026. The 5-minute studies used 126 UC San Diego psychology undergraduates and 158 Prolific workers, and the authors analysed 1,023 games. A win rate is the share of games in which the interrogator picked the machine as the human.

  • GPT-4.5 with the persona prompt: 73% (69% with undergraduates, 76% with Prolific workers)
  • LLaMa-3.1-405B with the persona prompt: 56% (45% with undergraduates, 65% with Prolific workers)
  • LLaMa-3.1-405B without the persona: 38%
  • GPT-4.5 without the persona: 36%
  • ELIZA, the 1960s rules-based chatbot used as a check: 23%
  • GPT-4o without the persona: 21%

The preprint called this the first evidence that any artificial system passes a standard three-party test. The published abstract is more cautious, saying the results are evidence that artificial systems can pass one. That changed wording is worth knowing before repeating a headline that says AI beat the Turing test.

What does passing a Turing test actually mean?

There is no official pass mark. In this paper a system passes when its win rate is not significantly different from 50%, the rate an interrogator would get by guessing. That is a null result, which is why the authors included ELIZA: a test that cannot catch a 1966 chatbot is not discriminating, and ELIZA scored significantly below chance. A win rate above 50% is stronger. GPT-4.5 with the persona was significantly above chance with both groups, meaning interrogators picked the machine over the real person more often than not.

Our arithmetic from the paper's table: GPT-4.5 with the persona won 185 of 254 games, so the real human was picked as the human in about 27% of its games. A 95% Wilson interval for 185 of 254 is about 67% to 78%, comfortably above 50%. For the GPT-5 follow-up (89 wins in 150 games) the same interval is about 51% to 67%, which only just clears 50%. The authors give a corrected p-value of 0.05 for that result. For the percentile arithmetic behind intervals like these, see our percentile calculator and the guide to IQ score margin of error.

The intervals are our own calculation from the win and loss counts printed in the paper, using the Wilson method. The paper reports its own significance tests, and its game counts differ slightly between the Methods text and Table 1 (1,021 against 1,023), so treat our intervals as approximate.

How much did the persona prompt matter?

A great deal. The persona prompt told the model to be a young person who is introverted, knows internet culture and uses slang; the authors chose it after an exploratory round that compared five prompts and seven models. Without it, GPT-4.5 fell from 73% to 36%, and the paper says neither GPT-4.5 nor GPT-4o without the persona did significantly better than ELIZA in either group. So the finding is about a model plus a prompt, not a model alone. The authors also note that the human witnesses were not told to role-play, and that humans might have matched the result if they had been told to play an introverted 19-year-old using slang.

The earlier two-party study makes the same point more sharply. Jones and Bergen's 2024 paper, presented at NAACL, tested 45 differently prompted GPT-3.5 and GPT-4 witnesses and reported that the best GPT-4 prompt was judged human in 49.7% of games, against 66% for real humans, 22% for ELIZA and 20% for the best GPT-3.5 prompt. The worst GPT-4 prompt scored 6%. The headline 49.7% is the best of many prompts, not the average GPT-4.

Did the 15-minute replication hold up?

Mostly, with a weaker margin. The PNAS version adds a third, preregistered study with 205 Prolific participants who played 396 games lasting 15 minutes each. GPT-4.5 had been withdrawn by then, so the authors tested GPT-5 with the same persona prompt, alongside LLaMa-3.1-405B and ELIZA. GPT-5 was judged human in 59% of games (89 wins, 61 losses), which the authors call marginally above chance after correction; LLaMa-3.1-405B reached 56%, not significantly above chance; ELIZA scored 17%. Interrogators identified the machine correctly in 44% and 41% of games, no better than chance. It also means the 73% GPT-4.5 result was never re-tested.

Two other cautions apply. The interrogators were lay people from a student pool and an online panel, and the authors suggest that expert interrogators, higher stakes or hour-long conversations might expose the models. And all three studies come from the same research group; we found no independent replication of the three-party result, nor any other human-judged Turing test with a new frontier model in 2026.

What did interrogators actually ask and notice?

This is the most useful finding for anyone who has wondered how to spot a chatbot. According to the paper, the most common strategy, in 61% of games, was small talk; in 50% the interrogator probed social and emotional qualities; in 19% they asked directly whether the witness was human; and in only 12% did they quiz knowledge or reasoning, including current events. The most frequent reasons for a verdict concerned linguistic style (27%) and interaction dynamics (23%).

One of the reasons most predictive of a correct verdict was that a witness lacked knowledge an AI should have: the model had to play dumb to seem human. The strategies that worked best were strange inputs and instruction-override attempts, and neither knowing about language models nor chatting with them often predicted accuracy. Turing himself imagined questions about chess, maths or poetry; the 2026 interrogators mostly did not ask them.

What did Turing actually predict, and what happened in 2014?

Turing's 1950 paper, Computing Machinery and Intelligence (Mind, volume 59, issue 236, pages 433 to 460), predicted that in about fifty years an average interrogator would have no better than a 70% chance of identifying the machine correctly after five minutes of questioning. The familiar 30% figure is the complement of that sentence, and it was a forecast, not a pass boundary. In June 2014 a chatbot posing as a 13-year-old Ukrainian boy, Eugene Goostman, was reported to have fooled 33% of judges at the Royal Society and was declared a pass by its organisers; Murray Shanahan of Imperial College London argued the next week that the 30% was Turing's prediction, not a threshold. The design differences matter: a persona built to excuse errors is not the three-party test run in 2026.

Is the Turing test an IQ test?

No, and the authors say so. The PNAS paper describes the test as not a direct test of intelligence but of humanlikeness, and adds that as machines have improved at maths and chess, intelligence alone is no longer enough to seem convincingly human. The 2024 paper reached the same view: the reasons interrogators gave rarely concerned knowledge or reasoning.

An IQ score is a different kind of number. It places a person on a distribution built from a norming sample, with a mean of 100, so it only has meaning relative to humans of a given age; a Turing result is a share of games in which a witness was judged human. We found no study that compares Turing-test judgments with IQ-style scores, and neither Jones and Bergen paper reports a psychometric score for any model; that comparison is our reasoning, not a cited finding. For what AI systems do score on reasoning tests, see our guides to whether AI can pass an IQ test, whether AI models have a g factor and why AI benchmark scores disagree, and the blog guides to the ChatGPT IQ score and what AGI means.

Passing a Turing test shows that a chatbot can be mistaken for a person in a short chat. It does not give the chatbot an IQ.

A June 2026 preprint by Jones and colleagues, Inverse Turing Bench, turns the question around and asks models to judge human against AI dialogue. Its abstract reports that GPTZero, Claude Opus-4.6 and GPT-5.5 reached 89.41%, 77.92% and 75.94% accuracy, and says statistical detection has semantic blind spots while semantic approaches are vulnerable to persona prompting. It is not peer reviewed, and it is about machine judges, not people.

Your own number

Where would your own score land?

Want a score on the standard IQ scale, with its percentile and a full report? Take the IQ Metrics test.

Find your IQ score now! →
Secure & encryptedInstant results10–20 minutes

To see what an IQ score means for a person, take the IQ Metrics test and read the result against the IQ bell curve. When a model, not a person, sits an assessment, other problems appear, which our guide to AI agents and proctored hiring tests covers, and our guide to the Cognitive Reflection Test shows one human reasoning test that chatbots now answer better than people.

Common questions

Has AI passed the Turing test?

In one peer-reviewed study, yes. A PNAS paper published on 19 May 2026 found GPT-4.5 with a humanlike persona prompt was judged human in 73% of five-minute, three-way chats. Without the persona it scored 36%, and a 15-minute test with GPT-5 gave 59%. All three studies are from the same group.

What percentage do you need to pass the Turing test?

There is no official pass mark. In the 2026 study a pass meant a win rate not significantly different from 50%. The 30% figure often quoted comes from Turing's 1950 prediction, which was a forecast and not a threshold.

Is the Turing test an IQ test?

No. The result is a share of games in which a witness was judged human, not a score on a normed scale. The authors describe the test as a measure of humanlikeness and say intelligence alone is not enough to appear convincingly human.

How long were the conversations in the Turing test study?

Five minutes in the two main studies, with 1,023 games analysed, and 15 minutes in a third replication with 396 games. In the five-minute games the median conversation lasted 4.2 minutes and 8 messages.

Sources for this story

  1. Large language models pass a standard three-party Turing test (Jones and Bergen), PNAS 123(21):e2524472123, published 19 May 2026 — PNAS
  2. Does GPT-4 pass the Turing test? (Jones and Bergen, 2024), NAACL 2024; arXiv 2310.20216 — Association for Computational Linguistics / arXiv
  3. Computing Machinery and Intelligence (A. M. Turing, 1950), Mind 59(236):433-460 — Mind, Oxford University Press
  4. Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue (Hager, Rathi, Hasan and Jones), arXiv 2606.21844, June 2026 — arXiv
  5. Professor disputes whether computer Eugene Goostman passed the Turing test (11 June 2014) — Imperial College London

Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.

Share this story

Know someone who keeps seeing this number quoted without the scale it was measured on? Send it to them — it takes one tap.

Filed under#ai benchmarks#study quality#scores and scales#test conditions#myths and fact checks

Read the research.
Then find your own number.

Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.

Start IQ Test →
Secure & encryptedInstant results10–20 minutes