Frontier models now operate a computer well enough to sit an assessment alongside a candidate. Across 19,368 AI-led interviews, 38.5 per cent were flagged for AI assistance.
Every standardised assessment rests on an assumption so basic it is rarely written down: that the person taking it is the person being measured, working alone, with only the permitted materials. Remote testing strained that assumption for a decade. In September 2026 a specific technical development broke it, and the assessment industry's own numbers show the break happened before the development arrived.
OpenAI released GPT-6 Astra on 3 September 2026, describing it as state of the art on computer use and browsing. The relevant figures are not the reasoning benchmarks that led the coverage but the ones measuring whether a system can operate an interface: 92.7 per cent on ScreenSpot-Pro, which tests locating controls in professional software, and 72.6 per cent on OSWorld 2.0, a suite of real desktop tasks. The company also reported a 1.9-times speed-up in task completion against its previous model.
It shipped simultaneously to consumer, business and enterprise plans, and through the OpenAI API, Microsoft Azure and Amazon Bedrock. Enterprise access is off by default and must be enabled by an administrator, which slows organisational adoption and does nothing at all to slow a candidate using a personal account on a second device.
The capability that matters here is narrow and unglamorous: reading what is on a screen, and acting on it, without a person driving. An assessment delivered through a browser is, from the software's perspective, just another interface.
This is not a forecast. Assessment vendors have been counting for over a year, and the figures are more striking than the speculation.
Detection is not the hard part. Of the candidates flagged for AI assistance, 61 per cent still cleared the bar.
That last figure is the one to hold on to. A flag is a signal, not a decision, and an organisation that flags more candidates than it can adjudicate has bought itself information it cannot act on. Meanwhile a screening test that lets through a majority of the people it has already identified as assisted is not performing the function it was bought for.
The instinct is to treat this as cheating, which frames it as a question of enforcement and character. The more useful frame is psychometric, and it is less comfortable.
A cognitive or aptitude test produces a raw score, and that score becomes interpretable only by comparison with a norming sample: a group of people who sat the same instrument under stated conditions. The published tables that convert a raw score into a percentile are a description of how that group performed under those conditions. Change the conditions and the tables no longer describe the new situation — not because the arithmetic breaks, but because the comparison group faced a different task.
This is the same principle behind a much less controversial rule: an untimed administration of a timed test yields a score that cannot be read against the timed norms. Nobody regards that as cheating. It is simply a different test. A remotely administered reasoning test in which some candidates have a capable assistant and others do not has the same problem, with the added difficulty that the administrator cannot tell which group any given candidate is in.
Cognitive ability tests remain among the better-validated tools in personnel selection, and the case for them was always narrower than either their advocates or their critics tend to state. We set out the evidence in our piece on what a cognitive ability test at work actually predicts: a real but moderate relationship with job performance, stronger in complex roles, and nothing like the deterministic filter the practice is often assumed to be.
That modest effect size is exactly why contamination matters so much. When an instrument explains a limited share of the variance in an outcome, a moderate amount of noise in the input can eliminate the signal. A screening test whose scores are partly a measure of who used an assistant is not a weaker predictor of performance. It is a predictor of something else.
Three responses are visible in the market, and they differ in how much they concede.
The third is the most interesting and the least proven. It is a genuine change of construct: no longer "how well does this person reason unaided" but "how well does this person direct a capable tool". Those are different abilities, they are presumably correlated, and how strongly is an empirical question nobody has answered yet. Any vendor claiming otherwise is ahead of the evidence.
Regulation is arriving behind all of this rather than in front of it. The European rules governing AI in hiring assessments, which had been expected this year, now land in 2027 — and they are largely concerned with the fairness of AI used by the employer, not with AI used by the candidate.
If you want to know how you actually reason, the conditions are the point. Sit a properly normed test once, unaided, and read the result as a position in a reference population with a range around it. A score produced any other way is measuring something, but not that.
Find your IQ score now! →The honest summary is that unsupervised remote cognitive screening was already under strain, and a model that can operate a computer competently has made the strain structural rather than incidental. The instruments are not broken; the delivery channel is. What employers do about that — better proctoring, live verification, or measuring a different thing on purpose — is unsettled, and worth watching more closely than the benchmark tables that prompted it.
Current models can read a screen and operate an interface well enough to assist in real time. GPT-6 Astra, released in September 2026, reports 92.7 per cent on ScreenSpot-Pro, which measures locating controls in professional software, and 72.6 per cent on the OSWorld desktop task suite. Assessment vendors have been recording AI-assisted attempts since 2025; one platform flagged 38.5 per cent of candidates across 19,368 interviews.
Partially. Vendors now describe four layers — identity verification, desktop restriction, in-test monitoring and post-submission analysis — on the explicit basis that no single control catches every tactic. Second devices and virtual machines that hide a separate AI session are specifically the routes browser-level controls cannot observe. Detection is also not the same as consequence: of candidates flagged on one platform, 61 per cent still scored above the passing bar.
The instruments themselves are unchanged, and supervised in-person administration is unaffected. What weakens is the interpretation of an unsupervised remote score, because a test's norm tables describe how a reference group performed under stated conditions. If those conditions cannot be guaranteed, the percentile attached to a score carries an assumption that may not hold for that candidate.
Three approaches are visible: layering more proctoring controls; adding a short live follow-up in which candidates explain their own submission, which is effective but costly per candidate; and restructuring the assessment to include an AI assistant deliberately and measure how well the candidate directs it. The third changes what is being measured, and how well it predicts job performance has not yet been established.
Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.
Aptitude tests are among the better-studied hiring tools, and for decades the headline validity figures were quoted with more confidence than the corrections behind them deserved. A 2022 reanalysis pulled those numbers down.
Recruitment screening and candidate assessment are high-risk uses under the EU AI Act, and their obligations were set to begin on 2 August 2026. A regulation in force since late July moved that date to 2 December 2027 — while leaving one workplace ban already biting.

A cognitive ability test can produce pass rates eight points apart across two groups and clear federal guidelines. Another test can produce an eleven-point gap and fail the same screen. The difference is a ratio, not a point spread — and it is worth calculating yourself.
Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.
Start IQ Test →