Skip to content

Product1 publisher3 min readPublished

A 33% detector score beat a Berkeley professor's argument. That is a publishing problem

Zvezdelina Stankova says she used AI only to edit her op-ed on Berkeley admissions. A Pangram reading of roughly 33% moved the story anyway, which is a byline problem, not a detection one.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying A 33% detector score beat a Berkeley professor's argument. That is a publishing problem
Photo: thenextweb.com

What happened

  • Zvezdelina Stankova published an op-ed on 15 August arguing that Berkeley admits students who cannot do middle school mathematics; within days the argument had been overtaken by a question about how the op-ed itself was written.
  • Stankova is a teaching professor of mathematics at UC Berkeley and made her case in the San Francisco Standard under the headline "I teach calculus at Berkeley. Some of my students can't do middle school math".
  • The piece was picked up by Fox News and Townhall and became a set-piece in the long fight over the University of California's test-blind admissions policy.
  • Pangram, the detector Substack uses to flag machine-written posts, was pointed at the op-ed and returned a reading of roughly 33% AI-generated or AI-assisted.
  • A post about the Pangram result by Chris Hoofnagle, a professor at Berkeley Law, drew close to three million views on X.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

On 15 August, Zvezdelina Stankova, a teaching professor of mathematics at UC Berkeley, published an op-ed in the San Francisco Standard under the headline "I teach calculus at Berkeley. Some of my students can't do middle school math" [1] [2]. Within days the argument had been overtaken by a question about how the op-ed itself was written [1], and the mechanism that did it is now available to anyone with a browser tab.

The piece travelled: Fox News and Townhall picked it up, and it became a set-piece in the long fight over the University of California's test-blind admissions policy [3]. Then Pangram, the detector Substack uses to flag machine-written posts, was pointed at it and returned a reading of roughly 33% AI-generated or AI-assisted [4]. A post about that result by Chris Hoofnagle, a professor at Berkeley Law, drew close to three million views on X [5]. Stankova told the Daily Californian she used AI to help edit the piece, and that the article represented "several hundred person-hours of intensive human work and deliberation, of which about 80 hours are my own" [6].

Note what she conceded and what she did not. Editing assistance is permissible under most campus policies, and the line between a tool that fixes sentences and one that writes them is precisely the line universities have spent three years failing to codify [7]. Hannes Bajohr, who teaches German at Berkeley and writes on machine authorship, told the same paper the episode "seems like deception", or at least something dishonest, absent a disclosure at the point of publication [8].

The detector reading did not have to be a finding to work as one. Pangram sits among the more credible tools in a thin field, but independent researchers have argued its false-positive rate is understated, and a 33% score is a probability estimate rather than a confession [9]. What no detector can currently do is distinguish a lightly edited human draft from a heavily prompted machine one, which is the case that actually matters [10]. So the score is not evidence of much. It is a headline, and in this instance it was enough.

The cost landed on the substance. Stankova reported that before 2020, when tests were still required, 71% of her Calculus I students were ready or nearly ready for the course, and that by 2023 the figure was 26%, with the most common diagnostic score being zero [11]. That is a 45-point drop in three years [12], sourced to her own classroom diagnostics, alongside admissions disparities between Bay Area schools and open letters signed by five Nobel laureates, among them Jennifer Doudna, calling for standardised tests to return [11] [13]. The university has published no rebuttal of either set of figures [13]. Those numbers are now harder to discuss than they were on 14 August.

The institutional vacuum is real. Berkeley's own computer science faculty have reported rising failure rates alongside heavier AI use in coursework [14], and under the EU's labelling regime an AI-written article can go unlabelled while a proofread email gets a marker [15]. American universities have nothing equivalent to point to at all [15]. That is an argument for writing your own policy, not for waiting.

For any team publishing under a named byline, the operational reading is narrow: a one-line disclosure of what assistance was used, applied before publication, is cheaper than arguing about a percentage afterwards. Neither Stankova nor the San Francisco Standard has said whether a disclosure will be added to the piece [16]. Watch whether the Standard adopts a standing policy rather than a one-off note, and whether any campus that disciplines students for AI use publishes the same rule for faculty bylines.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories