Skip to content

ScienceNot yet confirmed elsewhere1 publisher3 min readPublished

AI text detectors are good enough to deploy. The appeals process is what nobody has written.

Cancer-journal editors are already catching undisclosed LLM use in peer review with commercial software. The unresolved question is what happens to a flagged researcher who says no.

The Scientist · Science desk

How we use AISend a correction

What happened

  • The AACR screens peer-review reports with the commercial detector Pangram, over concerns that reviewers are using AI without the disclosure its policy requires.
  • Confronted by an AACR editor, one flagged reviewer admitted using a large language model after running out of time and complimented the detection.
  • Pangram Labs advertises 99.98% accuracy at detecting AI-generated content on its own website.
  • NeurIPS said in June that it rejected 18% of submissions after screening them with Pangram.
  • The tools still err, so their outputs can serve only as starting points for an investigation.

Why it matters

  • decision A research-integrity office with no written flag threshold and no named person to review flags has still chosen a policy, and will find out what it is under pressure from the first contested case.
  • constraint Because the software cannot separate mixed human-and-machine drafts, any policy leaning on it must define which uses are offences first; the score cannot supply the definition.
  • exposure Reviewers, authors and students are now scored by systems they cannot inspect, and in the absence of an appeals route a confession is the only clean way out of a flag.
  • contradiction The vendors' near-perfect accuracy claims and the analysts' starting-point-only caveat cannot both set an office's burden of proof; whichever it adopts decides who has to prove what.

The arithmetic in the marketing is worth doing before anyone drafts a policy. Pangram's website advertises 99.98% accuracy [2], which permits one error for every 5,000 documents screened [18]. GPTZero advertises 99% [8], which permits one in a hundred, fifty times as many [19]. Two products in the same market are not separated by a factor of fifty in quality. They are separated by what was measured and on whose text. Both figures are the firms' own, and Nature's reporting puts the independent view more narrowly: the tools call human-only text human almost all the time [3], and a flag is a starting point for investigation rather than a finding [13].

The AACR case shows what resolution looks like in practice. Daniel Evanko asked, and the reviewer confessed to using a model after running out of time [6]. He did not expect that, because most researchers do not disclose AI help [7]. Take the confession away and the editor is holding a probability and a denial, which is the part of the workflow nobody has designed. The person being scored cannot inspect the system doing the scoring, and proving you wrote your own sentences is not a task with a defined end.

The disputes will not mostly involve wholly synthetic text. Pangram's reading, in a study reported in January, put some AI-generated text in one in eight biomedical articles last year [10], and the tools are weakest precisely where that volume sits: writing in which human and machine prose intertwine, with no clean boundary for what counts as problematic use [14]. Tim Requarth of NYU Langone says the detectors are good for screening out places pumping out slop [15], and that case is easy. Deciding whether a model that tightened a reviewer's grammar breached a disclosure rule is a definitional question, and a score does not answer it.

Opting out is not on offer either. Pangram verdicts sit on arXiv papers through the alphaXiv mirror [11], on Substack posts since July [17], and on coursework at the University of Chicago [12]. Its co-founder Max Spero has publicly flagged journalists and, on one occasion, the Pope's social-media posts [16]. Complaints will therefore reach research-integrity offices from third parties whether or not the office screens anything itself. The accused will have a ready defence in the record: before this generation of tools, detection software was notoriously unreliable, in the account of Simon Fraser University computer scientist Marzena Karpinska [4], because the statistical approaches flagged human-authored text as AI too often [5]. An office that wants its flags to carry weight has to publish its threshold and its appeals route before it uses one in anger.

What to watch

  • Whether the AACR or another publisher publishes a screening threshold and an appeals route rather than handling flags case by case.
  • Whether NeurIPS discloses how much of its 18% rejection rate rested on the detector rather than other grounds.
  • Whether any false-positive rate measured by someone other than the vendors, on peer-review reports specifically, gets published.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence64
Adoption68
Hype gap+22
Incentives74
Confidence55
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Daniel Evanko, director of journal operations at the American Association for Cancer Research (AACR), and the AACR deploy the commercial AI-detection tool Pangram because of concerns over the number of peer-review reports submitted to their journals that appear to use AI without disclosure, contrary to the publisher's policy.

    ReportedSupportedSource: Nature, quoting Daniel Evanko of the AACR2 sources— create a free account to open themView cited source
  2. [2]

    Pangram Labs, the New York City start-up behind Pangram, states on its website: "Detect AI-generated content with 99.98% accuracy."

  3. [3]

    Independent analysts say the detectors do work in the sense that they correctly flag solely human-written content as human almost all the time, although no tool can be perfect.

Sources

1 independent publisher whose own reporting we read for this story.

  1. nature.com

    2 articles · August 24, 2026

    AI-detection tools have made huge leaps forward — how good are they?

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories