IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment

What Is ChatGPT’s IQ Score?

Scores & Scales

ChatGPT IQ Score: Why the Numbers Disagree

Search for ChatGPT’s IQ score and you will find 155, 136, 116 and under 100, all reported seriously, sometimes for the same model. They are not contradictions to be resolved. They are different tests, scored different ways, and the spread between them is the most informative thing about the whole exercise.

Chart of reported AI IQ figures from four different tests, ranging from below 100 on a culture-fair test to 155 on the verbal subtests of the Wechsler scale, showing how far the same systems spread across tests

There is no single ChatGPT IQ score. Published figures for the same family of systems range from slightly below 100 on one culture-fair test to 155 on the verbal half of a clinical instrument, with 116 and 136 reported in between. Every one of those numbers was produced honestly. They disagree because they come from different tests, administered under different conditions, and because none of them is an IQ in the sense the word has when a psychologist uses it.

That spread is worth understanding, because the same forces distort the score a person gets from an online test.

What IQ scores has ChatGPT actually been given?

Four results account for most of what circulates, and they are not measuring comparable things.

  • 155 on the Wechsler verbal subtests. A clinical psychologist administered parts of the Wechsler Adult Intelligence Scale and reported a verbal IQ of 155, above 99.9 percent of the American standardisation sample. Only the verbal subtests could be given: the performance subtests need eyes, ears and hands.
  • Under 100 on a culture-fair test. On the Worldwide IQ Test, a non-verbal assessment designed to minimise language and cultural loading, GPT-4o scored slightly below the population average of 100.
  • 116 and 136 on two other tests. OpenAI’s o3 model was reported at 116 on one assessment and at 136 on the public Mensa Norway test, which would place it above roughly 98 percent of people if it were a person.
  • Up to 151 on tracked weekly testing. The TrackingAI project has been administering the Mensa Norway test to frontier models every week; by September 2026 the top of that chart had reached 151.

A system cannot be simultaneously in the top 0.1 percent and below average. What varies is the test.

Why does the same system score 155 and under 100?

The Wechsler verbal subtests reward stored knowledge: vocabulary, general information, verbal similarities. That is the closest thing to a language model’s home ground, and the result reflects it. The culture-fair test does the opposite. It strips out language and prior knowledge deliberately and asks for pattern completion on abstract figures, which is precisely the ability these systems have found hardest.

In a person, this rarely happens, because human abilities correlate: someone with a top-percentile vocabulary usually does well on matrices too. In a machine there is no such tie between the two, so the choice of test decides the answer. The same principle explains why verbal and non-verbal index scores can diverge sharply in a human profile too, and why a single full-scale number can hide it.

Chart of reported AI IQ figures from four different tests, ranging from below 100 on a culture-fair test to 155 on the verbal subtests of the Wechsler scale, showing how far the same systems spread across tests
Chart of reported AI IQ figures from four different tests, ranging from below 100 on a culture-fair test to 155 on the verbal subtests of the Wechsler scale, showing how far the same systems spread across tests

How do you even give an IQ test to a chatbot?

Less straightforwardly than the headline numbers suggest, and the administration details do a lot of work. The Mensa Norway test behind most published AI figures is a set of 35 visual-pattern puzzles. A model that cannot see is given those puzzles described in words; a model that can see is given the original images. TrackingAI runs both variants weekly and reports the average of the last seven administrations rather than a single sitting.

That produces something genuinely instructive. The same underlying system often appears twice on the chart, once reading a description and once looking at the picture, and the two scores are not the same. One model scored 133 when the puzzles were described to it and 136 when it could see them. Another scored 113 reading and 108 looking. The direction is not even consistent.

  • The format of administration moved the score by several points without anything about the system changing.
  • Averaging seven sittings hides how much any single sitting varies, which for a person is exactly what a confidence interval is for.
  • A test built to be visual becomes a different test when it is read aloud, in the same way a timed test becomes a different test when the clock is removed.

Human testing treats this as fundamental rather than incidental. Standardised administration — same instructions, same time limit, same materials — is part of the instrument, because two scores only mean the same thing if they were produced the same way. Almost none of the AI figures in circulation were.

Your own number

Where would your own score land?

Take the IIF-certified assessment and get your score with the scale it was measured on, the percentile it corresponds to and the confidence range around it — the three figures most online tests leave out.

Find your IQ score now!

Secure & encryptedInstant results10–20 minutes

What happens when the test has never been published?

This is the most revealing comparison available, and it comes from the same project that produces the headline numbers. TrackingAI runs two tests on every model it follows. One is Mensa Norway, a public online test that has been on the internet for years and is therefore in the training data. The other is an offline test written by a Mensa member that, in the project’s own words, has never been on the public internet and is in no AI training data.

Reading the underlying chart data on 5 September 2026, 27 models carried a score on both tests. Across those 27, scores on the public test averaged about 13.5 points higher than scores on the unpublished one — close to a full standard deviation on the usual IQ scale. The public test topped out at 151; the unpublished test at 136.

  • One widely used model scored 137 on the public test and 113 on the unpublished one, a 24-point gap.
  • Another scored 145 publicly and 127 privately.
  • A third scored 108 publicly and 88 privately.
  • A small number moved the other way, which is a useful reminder that this is a pattern rather than a law.

An honest caveat belongs here, because it is the kind this site insists on. Two different tests are not directly comparable: they have different norming samples, different difficulty and different item formats, and the unpublished test’s construction has not been released. The gap is consistent with the public test having leaked into the training data, but it does not prove it on its own. What can be said plainly is that models score substantially higher on the test they have almost certainly seen.

The human parallel is exact, and it is the reason serious tests are kept out of circulation. Someone who has worked through a test before scores higher on it the second time without having become any smarter. That is the practice effect, and it is a measured, predictable thing rather than a suspicion.

Is an AI IQ score a real IQ?

No. An IQ is not a mark out of anything. It is a position within a reference population, expressed on a scale with a stated mean — almost always 100 — and a stated standard deviation, usually 15. Producing one requires a norming study in which the test is given to a representative sample of that population under standard conditions.

For a machine, there is no population to be a member of and no norming sample, so the final step simply cannot be taken. What a model produces on an IQ test is a count of correct answers. Calling that count an IQ borrows the authority of a scale it was never placed on. We make the full argument in our piece on why a model scoring 120 is not scoring an IQ.

There is a second problem specific to machines. Standard administration assumes a fixed time, no external help and no second attempt. A model may be run at different effort settings, with or without tools, once or many times, and the reported figure is usually the best of those runs. A human score reported that way would not be accepted either.

What this means for your own score

The lesson transfers directly. A number without a stated scale and a stated reference group is not a result, whoever produced it. When you take a test online, the questions that matter are which population your score is being compared against, what the scale’s standard deviation is, and how much measurement error sits around the figure — a point we cover in the explainer on why a score is a range, not a number.

It also means treating a score you got on a test you had already seen with the same scepticism you would apply to a model’s 151. If you want a number that means something, take a properly normed reasoning test once, cold, and read the result as a range rather than a point. For what the resulting figure does and does not tell you, how to read a test report is the place to start, and the wider human-versus-machine comparison puts the AI numbers in context.

Share this article

Know someone who keeps seeing these numbers quoted without the scale they were measured on? Send it to them.