Skip to content

Build1 publisher3 min readPublished

Repeating an AI classroom trial with 164 students erased the training advantage

Thibault Schrepel randomised his Vrije Universiteit Amsterdam law students into ban, unguided ChatGPT and trained arms, ran the design in 2024 and again in 2025, and the trained arm's lead was gone by the second year.

The Engineer · Build desk

Illustration accompanying Repeating an AI classroom trial with 164 students erased the training advantage

What happened

  • A law professor at Vrije Universiteit Amsterdam randomly split his "Law of AI" students into three arms: no ChatGPT, ChatGPT suggestions embedded in the text with no guidance, and hands-on training in prompting and checking output.
  • Every arm got the same task, with teams of four or five given 20 minutes to improve a provision of the EU AI Act, graded on substance, clarity, proportionality and innovation.
  • The trial ran with 66 students in 2024 and was repeated with 164 in 2025, using the same multiple-choice test and take-home revision exam both years.
  • The no-AI arm finished last in both years, and its teams regularly ran dry after ten to fifteen minutes, an effect the researcher calls "idea exhaustion."

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • contradiction Year one is the evidence for funding structured AI training and year two removes it, and Schrepel says he cannot tell whether the 2024 gain came from the training or from hands-on use.
  • constraint A timed team exercise and a graded take-home ranked the same students differently, so an enablement programme validated on a workshop task is not measuring what the graded work measures.
  • decision Once the measured training advantage is gone by the second cohort, the spending that still has a case behind it is faculty capability and assessment redesign.
  • precedent UC Berkeley Law's near-total ban on AI in graded work rests on the claim that core thinking skills must be built first, and this trial puts the burden of evidence on that claim.

The 2024 cohort was small. Sixty-six students split three ways is about 22 per arm [1], or roughly five subgroups per condition at teams of four or five [2]. The trained group's clearest lead was on the take-home exam [14], measured on that population. The 2025 repeat enrolled 164 students [7], about 55 per arm [3], and the three groups came out level [15]. Neither year's write-up gives scores or effect sizes [4].

The two instruments disagree about the same students. In the timed exercise, every subgroup in the unguided arm kept at least one misleading or legally extraneous term from ChatGPT's output [12], and some swapped "shall" and "individual" for AI-suggested alternatives without showing they understood the legal implications [11]. Schrepel expected that to carry into the exams, and it did not; the errors did not repeat, and the unguided group scored slightly higher than the AI-free group [19]. "I was wrong," he writes [20].

His explanation is that students learn to spot AI weaknesses through their own use, especially when accuracy has real consequences [21]. That is one instructor's reading of his own cohort. The two settings work differently: the exercise put ChatGPT's revisions directly inside the text under a 20-minute clock [2][4], and accepting a suggestion that is already in the document takes no keystrokes. The exams were a multiple-choice test and an individually graded take-home revision [6]. The arm with no tool at all did get something out of the constraint, which Schrepel records as deeper discussion among group members [9].

For the year-one result to transfer, your population has to look like the year-one population. Schrepel attributes the 2025 convergence to growing chatbot familiarity, and says many students already use these tools in daily life [16]. A team with thin baseline use is in the 2024 condition; a team where everyone already prompts daily is in the 2025 condition. The second transfer condition is the grader. Substance, clarity, proportionality and innovation [2] are judged by a person, and no test suite fails when a drafter keeps a legally extraneous term. Code review has a ground truth this task lacks, so the self-correction Schrepel observed may have a different cause.

Whether the advantage comes from structured training or just hands-on experience remains an open question, Schrepel says [18]. His policy conclusions do not wait for the answer: department leaders should resist blanket bans and give instructors room to experiment, and universities should invest in AI skills for faculty, many of whom feel poorly equipped [22]. He adds that ethical use and legal responsibility still have to be taught [25]. He also argues that pure literature reviews should no longer count toward a master's degree, and that programmes should require practice-oriented or empirical work in which thoughtful AI use is part of the grade [23].

What to watch

  • A third run with a fourth arm that gets hands-on use but no training would separate the two explanations Schrepel says he cannot yet distinguish.
  • Whether Vrije Universiteit Amsterdam actually changes master's requirements to stop counting pure literature reviews.
  • Whether UC Berkeley Law publishes outcome data from its near-total ban on AI in graded work.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories