IQ Metrics
IIF Certified Assessment Start IQ Test
IIF Certified Assessment
Work & Hiring TestsEmerging evidence

Hiring Tests Assumed the Candidate Was Alone. That Assumption Is Gone

Frontier models now operate a computer well enough to sit an assessment alongside a candidate. Across 19,368 AI-led interviews, 38.5 per cent were flagged for AI assistance.

Illustration generated for IQ Metrics. No photograph is used. IQ Metrics

What you need to know

  • GPT-6 Astra, released to business and enterprise customers in September 2026, reports 92.7 per cent on ScreenSpot-Pro and 72.6 per cent on OSWorld 2.0 — benchmarks of locating and operating controls in real software.
  • One assessment platform reviewing 19,368 AI-led interviews between July 2025 and January 2026 flagged 38.5 per cent of candidates for AI-assisted cheating, and 61 per cent of those flagged still scored above the passing bar.
  • Proctoring adoption among assessment customers rose from 64 per cent in January 2025 to 77 per cent by July, and the tactics it now has to catch include virtual machines that hide a second AI session from the proctoring software.
  • The psychometric consequence is specific: a test's norms describe performance under the conditions the norming sample faced. When those conditions can no longer be guaranteed, the score stops meaning what the table says.

Every standardised assessment rests on an assumption so basic it is rarely written down: that the person taking it is the person being measured, working alone, with only the permitted materials. Remote testing strained that assumption for a decade. In September 2026 a specific technical development broke it, and the assessment industry's own numbers show the break happened before the development arrived.

What changed in the software

OpenAI released GPT-6 Astra on 3 September 2026, describing it as state of the art on computer use and browsing. The relevant figures are not the reasoning benchmarks that led the coverage but the ones measuring whether a system can operate an interface: 92.7 per cent on ScreenSpot-Pro, which tests locating controls in professional software, and 72.6 per cent on OSWorld 2.0, a suite of real desktop tasks. The company also reported a 1.9-times speed-up in task completion against its previous model.

It shipped simultaneously to consumer, business and enterprise plans, and through the OpenAI API, Microsoft Azure and Amazon Bedrock. Enterprise access is off by default and must be enabled by an administrator, which slows organisational adoption and does nothing at all to slow a candidate using a personal account on a second device.

The capability that matters here is narrow and unglamorous: reading what is on a screen, and acting on it, without a person driving. An assessment delivered through a browser is, from the software's perspective, just another interface.

The measurements the industry already has

This is not a forecast. Assessment vendors have been counting for over a year, and the figures are more striking than the speculation.

  • Across 19,368 AI-led interviews run between July 2025 and January 2026, one platform flagged 38.5 per cent of candidates for AI-assisted cheating.
  • Of those flagged, 61 per cent still scored above the passing bar — meaning detection and consequence are different problems, and the second one is unsolved.
  • Proctoring adoption among that platform's assessment customers rose from 64 per cent in January 2025 to 77 per cent by July.
  • The tactics being counted include AI-generated code submissions, proxy candidates recruited through Discord and Telegram, off-camera assistance on a second device, and virtual machines that conceal a second AI session from the proctoring software entirely.

Detection is not the hard part. Of the candidates flagged for AI assistance, 61 per cent still cleared the bar.

That last figure is the one to hold on to. A flag is a signal, not a decision, and an organisation that flags more candidates than it can adjudicate has bought itself information it cannot act on. Meanwhile a screening test that lets through a majority of the people it has already identified as assisted is not performing the function it was bought for.

Why this is a measurement problem, not just a policy one

The instinct is to treat this as cheating, which frames it as a question of enforcement and character. The more useful frame is psychometric, and it is less comfortable.

A cognitive or aptitude test produces a raw score, and that score becomes interpretable only by comparison with a norming sample: a group of people who sat the same instrument under stated conditions. The published tables that convert a raw score into a percentile are a description of how that group performed under those conditions. Change the conditions and the tables no longer describe the new situation — not because the arithmetic breaks, but because the comparison group faced a different task.

This is the same principle behind a much less controversial rule: an untimed administration of a timed test yields a score that cannot be read against the timed norms. Nobody regards that as cheating. It is simply a different test. A remotely administered reasoning test in which some candidates have a capable assistant and others do not has the same problem, with the added difficulty that the administrator cannot tell which group any given candidate is in.

The consequence is worth stating precisely, because it is easy to overstate. This does not mean cognitive ability tests have stopped predicting anything. It means the specific reading — this candidate scored at the 84th percentile of the norm group — carries an assumption about test conditions that unsupervised remote delivery can no longer support. Supervised, in-person administration is unaffected.

What a score still predicts, and what it never did

Cognitive ability tests remain among the better-validated tools in personnel selection, and the case for them was always narrower than either their advocates or their critics tend to state. We set out the evidence in our piece on what a cognitive ability test at work actually predicts: a real but moderate relationship with job performance, stronger in complex roles, and nothing like the deterministic filter the practice is often assumed to be.

That modest effect size is exactly why contamination matters so much. When an instrument explains a limited share of the variance in an outcome, a moderate amount of noise in the input can eliminate the signal. A screening test whose scores are partly a measure of who used an assistant is not a weaker predictor of performance. It is a predictor of something else.

What employers are actually doing

Three responses are visible in the market, and they differ in how much they concede.

  • Layered proctoring. Vendors now describe four defensive layers — identity verification, desktop restriction, in-test monitoring and post-submission analysis — on the explicit basis that no single control catches every tactic. This is an arms race, and the second-device and virtual-machine routes are specifically the ones a browser-level control cannot see.
  • Verification by conversation. A short live follow-up in which a candidate explains their own submission is reported as one of the most effective available checks; someone who leaned on a model tends to come apart within a question or two. It is also expensive per candidate, which is what the automated screen existed to avoid.
  • Changing what is measured. At least one vendor has restructured its interview to embed an assistant deliberately and assess how well the candidate works with it, on the reasoning that if the tool cannot be excluded from the job it should not be excluded from the assessment.

The third is the most interesting and the least proven. It is a genuine change of construct: no longer "how well does this person reason unaided" but "how well does this person direct a capable tool". Those are different abilities, they are presumably correlated, and how strongly is an empirical question nobody has answered yet. Any vendor claiming otherwise is ahead of the evidence.

Regulation is arriving behind all of this rather than in front of it. The European rules governing AI in hiring assessments, which had been expected this year, now land in 2027 — and they are largely concerned with the fairness of AI used by the employer, not with AI used by the candidate.

Your own number

Where would your own score land?

If you want to know how you actually reason, the conditions are the point. Sit a properly normed test once, unaided, and read the result as a position in a reference population with a range around it. A score produced any other way is measuring something, but not that.

Find your IQ score now!
Secure & encryptedInstant results10–20 minutes

The honest summary is that unsupervised remote cognitive screening was already under strain, and a model that can operate a computer competently has made the strain structural rather than incidental. The instruments are not broken; the delivery channel is. What employers do about that — better proctoring, live verification, or measuring a different thing on purpose — is unsettled, and worth watching more closely than the benchmark tables that prompted it.

Common questions

Can AI cheat on an online hiring test?

Current models can read a screen and operate an interface well enough to assist in real time. GPT-6 Astra, released in September 2026, reports 92.7 per cent on ScreenSpot-Pro, which measures locating controls in professional software, and 72.6 per cent on the OSWorld desktop task suite. Assessment vendors have been recording AI-assisted attempts since 2025; one platform flagged 38.5 per cent of candidates across 19,368 interviews.

Does proctoring stop AI assistance?

Partially. Vendors now describe four layers — identity verification, desktop restriction, in-test monitoring and post-submission analysis — on the explicit basis that no single control catches every tactic. Second devices and virtual machines that hide a separate AI session are specifically the routes browser-level controls cannot observe. Detection is also not the same as consequence: of candidates flagged on one platform, 61 per cent still scored above the passing bar.

Are cognitive ability tests still valid for hiring?

The instruments themselves are unchanged, and supervised in-person administration is unaffected. What weakens is the interpretation of an unsupervised remote score, because a test's norm tables describe how a reference group performed under stated conditions. If those conditions cannot be guaranteed, the percentile attached to a score carries an assumption that may not hold for that candidate.

What are employers doing about AI-assisted candidates?

Three approaches are visible: layering more proctoring controls; adding a short live follow-up in which candidates explain their own submission, which is effective but costly per candidate; and restructuring the assessment to include an AI assistant deliberately and measure how well the candidate directs it. The third changes what is being measured, and how well it predicts job performance has not yet been established.

Sources for this story

  1. GPT-6 Astra: A new generation of intelligence, computer-use benchmarks and availability — OpenAI
  2. Reported figures on AI-assisted cheating rates and proctoring adoption across online technical assessments, 2025-2026 — HackerEarth
  3. Industry reporting on AI-era interview cheating tactics and detection layers — Hyring
  4. Schmidt and Hunter, and subsequent re-analyses, on the predictive validity of cognitive ability tests in personnel selection — Psychological Bulletin
  5. Standards for Educational and Psychological Testing, on norming samples and conditions of administration — American Educational Research Association, American Psychological Association and National Council on Measurement in Education

Corrections: spotted an error? Email corrections@iqmetrics.org and we will update this story and note the change here.

Share this story

Know someone who keeps seeing this number quoted without the scale it was measured on? Send it to them — it takes one tap.

Filed under#work and hiring#test conditions#standardized exams#ai benchmarks

Read the research.
Then find your own number.

Our IIF-certified assessment reports your score with its scale, percentile and confidence range — and a breakdown of the cognitive domains behind it.

Start IQ Test
Secure & encryptedInstant results10–20 minutes