Science1 distinct publisher3 min readPublished
A HEPI survey puts generative-AI use on assessed work at 94% of UK undergraduates, and the educators Nature canvassed have largely stopped hunting for it, setting tasks instead where a model's output does not fit the evidence in the room.
The Scientist · Science desk
Compiled by The ScientistSomething wrong?How this is made
The engineering common to these redesigns is narrow and worth naming: each ties the assessed output to something the model cannot see. Ruben Verborgh's web-development exam at Ghent University is open book and open internet, on the reasoning that AI is part of the internet, and asks why a particular site is broken or slow, questions with several defensible answers that chatbots tend to answer without sound methodology [9]. Etienne Roesch, at the University of Reading, hands students screenshots of a specific analysis of neurological-activity data and asks for a narrative about it, so verbatim chatbot output often describes something else entirely [10]. Daniel Silver, at the University of Toronto Scarborough, replaced an essay on Adam Smith's 1776 book with a task in which AI agents trade in a simulated market and students must quote the transcripts their own agents produced [8]. The failure mode of an unchecked model answer surfaces in the marking, not in a misconduct hearing.
The survey arithmetic supports that reading better than the headline figure does. Roughly 94% of 1,054 respondents is about 990 students using the tools on assessed work [1], while the 12% who said they inserted AI-generated text directly into coursework is about 127 [2]. The 82-point gap, some 863 respondents, is where the word "help" does its work [3], and the instrument does not say whether that means summarising a reading or writing the paragraph that gets marked [1][2].
The American number is not a rival estimate. Nine percent of more than 95,000 students is at least 8,550 people [4], and the question asked about using AI on coursework while knowing it broke the rules [3], a condition the UK item did not impose. What the pair does not tell you is whether student behaviour differs across the two systems at all.
Detection signals carry their own selection problem. Hallucinated references and AI-typical punctuation [4], including the stretch Nikita Bezrukov describes in which half the cited literature did not exist [5], identify students who did not check the output, which is a sample skewed toward the careless. One response is to make the checking itself the assessed skill: Mattei's history-of-technology students run three LLMs over sources they have found and earn extra points for catching hallucinated references or infographics that misstate them [7].
As evidence, the Nature piece is a set of practitioner accounts gathered by its careers team [12]. It reports no comparison of marks or learning between the redesigned tasks and the assessments they replaced, and no estimate of how much misuse each design deflects [5]. Mattei's line about nobody having a right answer is the honest summary [6]. The cost side, though, is legible without a trial: someone has to read the transcripts and the interaction logs [11], and that someone is the marker. In a seminar that is a workload change. In a large service course it becomes a staffing question, and that is the constraint likely to decide how far this style of redesign travels.
Ranked by verification strength, evidence, and original report placement.
A 2026 survey of 1,054 UK undergraduate students by the Higher Education Policy Institute in Oxford found that roughly 94% said they use generative AI tools to help them with assessed work.
In the same HEPI survey, 12% of respondents said they directly inserted AI-generated text into their coursework.
A study published in May, using survey data from more than 95,000 students at 20 US universities, estimated that 9% of students used AI on their coursework despite knowing that this broke the rules.
Educators worldwide report suspicious signs of AI use, from hallucinated references to punctuation styles typical of AI tools, in take-home coursework and online examinations, fuelling concern about unfair advantage.
Nikita Bezrukov, who teaches linguistics and communication at MIT, says that at one point people were submitting essays in which "50% of the reference literature did not exist in real life".
Computer scientist Nicholas Mattei at Tulane University says "I don't think anybody has a 'right' answer right now" and that "things are changing pretty rapidly".
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
science
An LLM asked only to fix grammar stripped the personality markers out of human prose1 distinct publisher
build
Thirty runs of one prompt put a 34-point error bar around a brand's mention rate1 distinct publisher
invest
A Connecticut judge just priced prompt injection: no fine, no e-filing2 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Two surveys, seven classrooms
The quantitative spine is solid but borrowed: 94% and 12% from HEPI's 1,054 respondents, 9% from a May study of more than 95,000 US students, both reaching readers through Nature's summary of someone else's fieldwork. Everything after that is testimony from academics about their own teaching, named and specific — Silver's Adam Smith agents, Roesch's screenshots — and entirely unaudited.
Behaviour saturated, remedy scattered
Two different things are being adopted at very different rates. Student AI use on assessed work is effectively universal in the UK sample and confessed as rule-breaking by nearly one in ten US students. The counter-move is seven instructors on four continents changing their own syllabi; nothing in this reporting tells you a single department, faculty or exam board has standardised on any of it.
Slightly ahead of its proof
The promise implied by the framing — a rubric a chatbot cannot pass — is vouched for only by the people who wrote the rubrics. Verborgh says chatbots lack sound methodology; Roesch says the mismatch "shows"; Silver says shortcut submissions were "easy to spot". Plausible, all of it, and untested. Nature keeps the overshoot small by quoting Mattei up front that nobody has the right answer yet.
Instructors describing their own courses
Each fix is narrated by the person who built it and grades it, which is not disqualifying but is a direction of pull. The prevalence numbers come from a higher-education policy institute whose business is the sector's debates. Absent from the room: students, examination boards, and the detection-software industry whose product these academics have largely given up on.
One outlet, one method
The survey arithmetic would survive checking. The interesting part — that a well-built task makes chatbot-only work fail on its own terms — rests on a single publisher's canvass and on the word of the seven people who set the tasks. Enough to act on as practice; not enough to cite as a finding.