Machines have been writing reasoning items since long before the current wave of language models, and the research is clear about where the difficulty lies. Generating a plausible question is the easy half. Knowing how hard it is remains the expensive half.
Ask a current language model to write you a matrix reasoning puzzle and it will produce something that looks exactly right: a three-by-three grid, a missing cell, eight plausible options and a confident explanation of the rule. The obvious question follows immediately. If questions can be produced on demand at no cost, what happens to the tests built from them, and what happens to a score? The answer is more settled than the novelty suggests, because psychometricians have been generating items by machine for about twenty years and know precisely which part of the job is hard.
The field calls it automatic item generation, and the principle predates the current wave entirely. You do not write items; you write an item model — a template with defined slots and a set of rules for filling them — and the computer enumerates the variants. For figural reasoning this works particularly well, because the content of a matrix item is genuinely rule-governed: a shape rotates, an element is added, a count increments, and difficulty rises with the number and type of rules combined in one item. Generators of this kind for figural matrices and endless-loop items have been built and Rasch-calibrated in the published literature for two decades, and Gierl and Lai's instructional module for the National Council on Measurement in Education is the standard treatment of the method.
A concrete example: in 2018 Blum and Holling published a study in Frontiers in Psychology presenting the IMak package, an open tool for generating figural analogies. They produced 23 items and administered them to 307 participants. The results supported the exercise — the generated items showed adequate psychometric properties and convergent validity, most of the manipulated rules contributed to difficulty as intended, and difficulty could be predicted from the generating rules to some extent. That qualifier is the interesting part, and it is doing a great deal of work.
Writing the question was never the bottleneck. Knowing how hard the question is, before a single person has answered it, is the bottleneck.
A number from a test means something only because of machinery sitting behind it that has nothing to do with how the questions were written. Three pieces in particular:
None of the three can be generated. Each is an empirical fact about how real people responded, and each costs time and a sample to obtain. This is why an infinite supply of free questions does not produce a test: it produces an item pool, which is the cheap raw material a test is made from. Our checklist on what to look for before trusting an online score is essentially a list of whether this machinery exists behind the questions you were shown.
The older rule-based generators and the current language models fail in opposite directions, and the contrast is instructive. A rule-based generator is constrained by construction: it can only produce items its rules permit, so it knows what it made, but its range is narrow. A language model has no such constraint, which is exactly the problem — it produces fluent, varied, superficially excellent items with no guarantee that any of them measure the thing they appear to measure.
That is the pattern the recent research keeps finding. Reviews of language-model item generation report that models can produce items in vast quantity while many of them align only loosely with the construct being assessed, and that semantic drift can yield items with obvious face validity — they look like proper questions to a human reader — that nonetheless fail statistical quality criteria when administered. Work examining a situational judgement test written by an earlier model found the first version fitting the measurement model poorly, with an acceptable version arriving only after items were cut and the instrument refined against real response data. The refinement, not the generation, was where the quality came from.
For anyone taking a test, the practical consequence is that the origin of the questions is not the thing to ask about. A well-normed instrument built from machine-generated items is sound. A poorly built instrument written by hand is not. The distinguishing question is unchanged and it is about evidence, not authorship:
There is also a claim worth separating out, because it gets conflated with this one constantly. A machine writing test questions is a different thing from a machine answering them, and a model's performance on a reasoning benchmark is not an IQ in any sense the term supports — the argument is set out in full in our piece on why an AI scoring on an IQ-style test is not an IQ. Generation and performance are separate questions and neither settles the other.
Generated items are fine. Ungrounded scores are not. Take an assessment that names its scale, reports a percentile and tells you the range around your result.
Find your IQ score now! →The likeliest near-term effect of cheap item generation is not a flood of new tests but a change in how existing ones are maintained. Large calibrated item pools make it far easier to give two people different questions of matched difficulty, which reduces the value of memorising items and blunts the practice effect we describe in our piece on what a retest actually measures. That is a real improvement, and it is unglamorous, and it arrives only for tests that did the empirical work anyway. If you want a number from an instrument that did it, you can take the assessment and read the percentile it returns. The questions were never the scarce resource. The evidence about the questions always was, and no amount of generation produces it.
It can write plausible ones, and rule-based generators have produced psychometrically adequate ones for years – a 2018 Frontiers in Psychology study generated 23 figural analogy items with the IMak package and found adequate properties in 307 participants. But validity is not a property of the question alone. An item becomes usable only after its difficulty is estimated on a real sample, and research on language-model-generated items repeatedly finds items that look convincing while failing statistical quality checks.
That depends entirely on whether the test was normed, not on who or what wrote the items. A score is interpretable only against a reference sample that fixes the mean and standard deviation, and against a measured reliability that determines the confidence range around the result. A test built from generated items with all of that in place is sound; one without it produces a number that means nothing, regardless of how good the questions look.
It is an established psychometric method in which a test developer writes an item model – a template with defined slots and rules for filling them – and a computer enumerates the variants. It has been applied to figural matrices and analogy items for around two decades, with published work on generating Rasch-calibrated items, and it predates the current wave of large language models by a long way.
Not on that basis. Ask instead whether the test names the scale its score is on, describes the reference sample the score is compared against, reports a reliability figure or confidence interval, and gives a percentile alongside the number. Those questions separate a usable test from an unusable one far more reliably than knowing how the items were authored.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
IQ scales are defined by a human reference sample, and a language model is not in it. Placing a machine on that scale is not a hard measurement problem — it is a category error.
Two people can perform identically on two different online tests and walk away with scores twenty points apart. The scale, the comparison group and the margin of error are what separate an assessment from a quiz.
GPT-6 Astra scored 62.7 per cent and 99.9 per cent on the same benchmark in the same week. The difference is not the model. It is how the test was administered.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →