Science1 publisher3 min readPublished
AI-written answers to Wollongong law exams now outscore most of the students who sat them
Wollongong law researchers found nine AI models averaged 76.3% on a criminal law exam, beating 82.5% of students, up from 52.5% in their 2023 test. The sample is small, 18 scripts at one university, but the authors say it already calls into question grades from unsupervised assessments.
The Scientist · Science desk

What happened
- In tort law the AI papers averaged 66%, ahead of 61% of the students who sat the same exam.
- Seven of the 18 AI-written papers ranked at or above the 90th percentile of the student cohort.
- On hypothetical legal scenarios in criminal law, AI averaged 76% against 59.6% for students, while in torts the two were roughly level.
- Tutors blind-graded 12 of the AI papers mixed in with genuine student scripts, and the authors graded the other six themselves.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- exposure Grades from take-home and other unsupervised law assessments are now open to doubt, since subscription tools can produce scripts that outscore most of a cohort.
- constraint A high exam mark cannot be taken as evidence that a model is safe for legal research, because the same answers mixed strong analysis with fabricated authorities.
- decision Universities that want graduates fluent with AI still have to find an assessment format that tests whether students can catch its errors on their own.
The criminal law average rose 23.8 percentage points [1]. The rank moved further than the mark, from around the 22nd percentile in 2023 [9] to ahead of 82.5% of the cohort [6], a climb of roughly 60 percentile points [2]. The team's 2023 study had concluded that generative AI was not close to replacing humans in intellectually demanding tasks such as an undergraduate law exam [1]. Its specific finding was weak critical analysis of hypothetical legal scenarios. This time that weakness had largely disappeared [17].
Part of the gain may belong to the set-up. The models answered with internet access, and some used an enhanced reasoning mode that gives them more time to plan and check before answering [4]. They got no lecture notes, textbooks or curated legal materials [4], so whatever law they used came from training or the open web.
The design choice I admire is the blind marking. The dozen scripts slipped into the student pile [5] are the closest thing the study has to a control: the same tutors, the same guidelines, no knowledge that a machine wrote them. The six marked by the subject coordinators, who are also the authors, were not blind [5]. The published summary does not split scores by grader, so a reader cannot check whether knowing the source moved the marks.
The denominator is small. Every subject average rests on nine scripts from one university [3][14]. The authors call the capabilities jagged: performance varied substantially between models and subjects, and some models were excellent in one subject and weak in the other [12]. They adjusted prompts to keep answers a realistic length [18]. They did not test whether results hold across prompts, model versions or assessment formats [19], and they note that grading involves subjective judgement despite the guidelines tutors received [20].
The thing this doesn't tell you is whether the answers are reliable. According to the authors, strong analysis could sit alongside poor citations, weak source selection or fabricated authorities [13]. Some models hallucinated less, but they also cited fewer sources [13]. I think part of that lower rate is caution: a model that cites fewer cases has fewer chances to invent one. An exam mark measures legal analysis. In legal work the error that does damage is the invented case, and a good mark does not rule one out.
The authors' conclusion for universities is narrow, and in my view the evidence supports it. Most of the tools they tested are available by subscription [15]. The authors say the models are now capable enough that students could cheat with them, and that this calls into question results from unsupervised assessments [11]. They also argue that students still need enough foundational knowledge and skills to detect errors [16]. The fabricated authorities show why. A student who hands an unsupervised paper to one of these models can land in the top tenth of the cohort [8] without ever practising how to spot an invented case.
What to watch
- Replications at other universities or in other subjects such as contracts, which the authors flag as untested.
- Tests of the same models under varied prompts and newer versions, the consistency check this study did not run.
- Scores split between the 12 blind-graded papers and the six graded by the authors.