IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment
AI & Machine IntelligenceEmerging evidence

AI-Generated IQ Test Questions: What Changes for a Score

Machines have been writing reasoning items since long before the current wave of language models, and the research is clear about where the difficulty lies. Generating a plausible question is the easy half. Knowing how hard it is remains the expensive half.

Illustration generated for IQ Metrics. No photograph is used. IQ Metrics

What you need to know

  • Automatic item generation is an established branch of psychometrics, not a product of the current language-model wave. Rule-based generators for figural matrices and analogies have been built and calibrated for around two decades.
  • A 2018 study in Frontiers in Psychology generated 23 figural analogy items with the open IMak package and administered them to 307 people, reporting adequate psychometric properties and difficulty that was predictable from the generating rules to a limited extent.
  • Research on language-model-generated items has repeatedly found the opposite failure: items that read convincingly to a human reviewer while failing the statistical quality checks, a mismatch described in the literature as face validity without measurement quality.
  • None of this changes what makes a score interpretable. An item is only worth answering once its difficulty has been estimated on a real sample, and a score is only meaningful against norms collected from real people.

Ask a current language model to write you a matrix reasoning puzzle and it will produce something that looks exactly right: a three-by-three grid, a missing cell, eight plausible options and a confident explanation of the rule. The obvious question follows immediately. If questions can be produced on demand at no cost, what happens to the tests built from them, and what happens to a score? The answer is more settled than the novelty suggests, because psychometricians have been generating items by machine for about twenty years and know precisely which part of the job is hard.

Automatic item generation is not new

The field calls it automatic item generation, and the principle predates the current wave entirely. You do not write items; you write an item model — a template with defined slots and a set of rules for filling them — and the computer enumerates the variants. For figural reasoning this works particularly well, because the content of a matrix item is genuinely rule-governed: a shape rotates, an element is added, a count increments, and difficulty rises with the number and type of rules combined in one item. Generators of this kind for figural matrices and endless-loop items have been built and Rasch-calibrated in the published literature for two decades, and Gierl and Lai's instructional module for the National Council on Measurement in Education is the standard treatment of the method.

A concrete example: in 2018 Blum and Holling published a study in Frontiers in Psychology presenting the IMak package, an open tool for generating figural analogies. They produced 23 items and administered them to 307 participants. The results supported the exercise — the generated items showed adequate psychometric properties and convergent validity, most of the manipulated rules contributed to difficulty as intended, and difficulty could be predicted from the generating rules to some extent. That qualifier is the interesting part, and it is doing a great deal of work.

Writing the question was never the bottleneck. Knowing how hard the question is, before a single person has answered it, is the bottleneck.

What a score actually requires

A number from a test means something only because of machinery sitting behind it that has nothing to do with how the questions were written. Three pieces in particular:

  • Calibration. Each item needs an estimated difficulty, obtained by giving it to people and observing who answers it correctly. A generator can predict difficulty from its rules approximately; it cannot know it.
  • Norms. A raw count of correct answers becomes an IQ score only by comparison with a reference sample of known composition, which fixes the mean at 100 and the spread by a chosen standard deviation. Without that sample there is no scale for the score to sit on.
  • Reliability. The test needs a measured consistency, which determines the standard error attached to every score it produces — the reason a result is a band rather than a point, as our piece on the margin of error on an IQ score sets out.

None of the three can be generated. Each is an empirical fact about how real people responded, and each costs time and a sample to obtain. This is why an infinite supply of free questions does not produce a test: it produces an item pool, which is the cheap raw material a test is made from. Our checklist on what to look for before trusting an online score is essentially a list of whether this machinery exists behind the questions you were shown.

Where language models specifically go wrong

The older rule-based generators and the current language models fail in opposite directions, and the contrast is instructive. A rule-based generator is constrained by construction: it can only produce items its rules permit, so it knows what it made, but its range is narrow. A language model has no such constraint, which is exactly the problem — it produces fluent, varied, superficially excellent items with no guarantee that any of them measure the thing they appear to measure.

That is the pattern the recent research keeps finding. Reviews of language-model item generation report that models can produce items in vast quantity while many of them align only loosely with the construct being assessed, and that semantic drift can yield items with obvious face validity — they look like proper questions to a human reader — that nonetheless fail statistical quality criteria when administered. Work examining a situational judgement test written by an earlier model found the first version fitting the measurement model poorly, with an acceptable version arriving only after items were cut and the instrument refined against real response data. The refinement, not the generation, was where the quality came from.

There is a second problem specific to figural reasoning that plausibility hides. A well-formed matrix item has exactly one defensible answer. A generated item can easily admit two — a second rule that also completes the grid, which the generator did not consider and a reviewer skimming the options will not notice. Items like this look fine and behave badly, discriminating poorly because capable candidates find the alternative rule. Our walk-through of the one rule you should be able to say out loud is the reader-side version of the same test.

What this changes for you

For anyone taking a test, the practical consequence is that the origin of the questions is not the thing to ask about. A well-normed instrument built from machine-generated items is sound. A poorly built instrument written by hand is not. The distinguishing question is unchanged and it is about evidence, not authorship:

  • Is the scale named — the mean and standard deviation the score is expressed on? A number without them cannot be interpreted.
  • Is there a described reference sample the score is compared against, with some account of its size and composition?
  • Is a reliability figure or a confidence interval reported alongside the result, rather than a bare number?
  • Is the result given as a percentile as well as a score, so it survives translation between scales? Our percentile calculator does that conversion for any score you already hold.
  • Does the test explain what it measures, rather than promising to measure intelligence in general?

There is also a claim worth separating out, because it gets conflated with this one constantly. A machine writing test questions is a different thing from a machine answering them, and a model's performance on a reasoning benchmark is not an IQ in any sense the term supports — the argument is set out in full in our piece on why an AI scoring on an IQ-style test is not an IQ. Generation and performance are separate questions and neither settles the other.

Your own number

Where would your own score land?

Generated items are fine. Ungrounded scores are not. Take an assessment that names its scale, reports a percentile and tells you the range around your result.

Find your IQ score now!
Secure & encryptedInstant results10–20 minutes

The likeliest near-term effect of cheap item generation is not a flood of new tests but a change in how existing ones are maintained. Large calibrated item pools make it far easier to give two people different questions of matched difficulty, which reduces the value of memorising items and blunts the practice effect we describe in our piece on what a retest actually measures. That is a real improvement, and it is unglamorous, and it arrives only for tests that did the empirical work anyway. If you want a number from an instrument that did it, you can take the assessment and read the percentile it returns. The questions were never the scarce resource. The evidence about the questions always was, and no amount of generation produces it.

Common questions

Can AI write valid IQ test questions?

It can write plausible ones, and rule-based generators have produced psychometrically adequate ones for years – a 2018 Frontiers in Psychology study generated 23 figural analogy items with the IMak package and found adequate properties in 307 participants. But validity is not a property of the question alone. An item becomes usable only after its difficulty is estimated on a real sample, and research on language-model-generated items repeatedly finds items that look convincing while failing statistical quality checks.

Are AI-generated IQ tests accurate?

That depends entirely on whether the test was normed, not on who or what wrote the items. A score is interpretable only against a reference sample that fixes the mean and standard deviation, and against a measured reliability that determines the confidence range around the result. A test built from generated items with all of that in place is sound; one without it produces a number that means nothing, regardless of how good the questions look.

What is automatic item generation?

It is an established psychometric method in which a test developer writes an item model – a template with defined slots and rules for filling them – and a computer enumerates the variants. It has been applied to figural matrices and analogy items for around two decades, with published work on generating Rasch-calibrated items, and it predates the current wave of large language models by a long way.

Should I avoid a test that uses AI-generated questions?

Not on that basis. Ask instead whether the test names the scale its score is on, describes the reference sample the score is compared against, reports a reliability figure or confidence interval, and gives a percentile alongside the number. Those questions separate a usable test from an unusable one far more reliably than knowing how the items were authored.

Sources for this story

  1. Blum and Holling, "Automatic Generation of Figural Analogies With the IMak Package" (2018) — Frontiers in Psychology
  2. Gierl and Lai, "Using Automated Processes to Generate Test Items", ITEMS instructional module — National Council on Measurement in Education
  3. Mini-review on the impacts of artificial intelligence on the development of measurement scales (2026) — Frontiers in Organizational Psychology
  4. Research on evaluating the instrumental quality of LLM-generated assessment items (2026) — Frontiers in Education
  5. Freund, Hofer and Holling, "Explaining and Controlling for the Psychometric Properties of Computer-Generated Figural Matrix Items" (2008) — Applied Psychological Measurement

Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.

Share this story

Know someone who keeps seeing this number quoted without the scale it was measured on? Send it to them — it takes one tap.

Filed under#ai benchmarks#study quality#raven matrices#percentiles and norms#scores and scales

Read the research.
Then find your own number.

Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.

Start IQ Test
Secure & encryptedInstant results10–20 minutes