ScienceNot yet confirmed elsewhere1 publisher3 min readPublished
AI text detectors are good enough to deploy. The appeals process is what nobody has written.
Cancer-journal editors are already catching undisclosed LLM use in peer review with commercial software. The unresolved question is what happens to a flagged researcher who says no.
The Scientist · Science desk
What happened
- The AACR screens peer-review reports with the commercial detector Pangram, over concerns that reviewers are using AI without the disclosure its policy requires.
- Confronted by an AACR editor, one flagged reviewer admitted using a large language model after running out of time and complimented the detection.
- Pangram Labs advertises 99.98% accuracy at detecting AI-generated content on its own website.
- NeurIPS said in June that it rejected 18% of submissions after screening them with Pangram.
- The tools still err, so their outputs can serve only as starting points for an investigation.
Why it matters
- decision A research-integrity office with no written flag threshold and no named person to review flags has still chosen a policy, and will find out what it is under pressure from the first contested case.
- constraint Because the software cannot separate mixed human-and-machine drafts, any policy leaning on it must define which uses are offences first; the score cannot supply the definition.
- exposure Reviewers, authors and students are now scored by systems they cannot inspect, and in the absence of an appeals route a confession is the only clean way out of a flag.
- contradiction The vendors' near-perfect accuracy claims and the analysts' starting-point-only caveat cannot both set an office's burden of proof; whichever it adopts decides who has to prove what.
The arithmetic in the marketing is worth doing before anyone drafts a policy. Pangram's website advertises 99.98% accuracy [2], which permits one error for every 5,000 documents screened [18]. GPTZero advertises 99% [8], which permits one in a hundred, fifty times as many [19]. Two products in the same market are not separated by a factor of fifty in quality. They are separated by what was measured and on whose text. Both figures are the firms' own, and Nature's reporting puts the independent view more narrowly: the tools call human-only text human almost all the time [3], and a flag is a starting point for investigation rather than a finding [13].
The AACR case shows what resolution looks like in practice. Daniel Evanko asked, and the reviewer confessed to using a model after running out of time [6]. He did not expect that, because most researchers do not disclose AI help [7]. Take the confession away and the editor is holding a probability and a denial, which is the part of the workflow nobody has designed. The person being scored cannot inspect the system doing the scoring, and proving you wrote your own sentences is not a task with a defined end.
The disputes will not mostly involve wholly synthetic text. Pangram's reading, in a study reported in January, put some AI-generated text in one in eight biomedical articles last year [10], and the tools are weakest precisely where that volume sits: writing in which human and machine prose intertwine, with no clean boundary for what counts as problematic use [14]. Tim Requarth of NYU Langone says the detectors are good for screening out places pumping out slop [15], and that case is easy. Deciding whether a model that tightened a reviewer's grammar breached a disclosure rule is a definitional question, and a score does not answer it.
Opting out is not on offer either. Pangram verdicts sit on arXiv papers through the alphaXiv mirror [11], on Substack posts since July [17], and on coursework at the University of Chicago [12]. Its co-founder Max Spero has publicly flagged journalists and, on one occasion, the Pope's social-media posts [16]. Complaints will therefore reach research-integrity offices from third parties whether or not the office screens anything itself. The accused will have a ready defence in the record: before this generation of tools, detection software was notoriously unreliable, in the account of Simon Fraser University computer scientist Marzena Karpinska [4], because the statistical approaches flagged human-authored text as AI too often [5]. An office that wants its flags to carry weight has to publish its threshold and its appeals route before it uses one in anger.
What to watch
- Whether the AACR or another publisher publishes a screening threshold and an appeals route rather than handling flags case by case.
- Whether NeurIPS discloses how much of its 18% rejection rate rested on the detector rather than other grounds.
- Whether any false-positive rate measured by someone other than the vendors, on peer-review reports specifically, gets published.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence64
- Adoption68
- Hype gap+22
- Incentives74
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Daniel Evanko, director of journal operations at the American Association for Cancer Research (AACR), and the AACR deploy the commercial AI-detection tool Pangram because of concerns over the number of peer-review reports submitted to their journals that appear to use AI without disclosure, contrary to the publisher's policy.
ReportedSupportedSource: Nature, quoting Daniel Evanko of the AACR2 sources— create a free account to open themView cited source - [2]
Pangram Labs, the New York City start-up behind Pangram, states on its website: "Detect AI-generated content with 99.98% accuracy."
- [3]
Independent analysts say the detectors do work in the sense that they correctly flag solely human-written content as human almost all the time, although no tool can be perfect.
- [4]
Until Pangram and others came along, AI-detection software was notoriously unreliable, says Marzena Karpinska, a computer scientist at Simon Fraser University who has conducted independent analyses of the tools.
ReportedSupportedSource: Marzena Karpinska, Simon Fraser University2 sources— create a free account to open themView cited source - [5]
Earlier statistical approaches, based on perplexity, burstiness and stylistic quirks such as over-use of the 'It's not X, it's Y' construction, too often falsely flagged human-authored text as AI-generated.
- [6]
When Evanko asked a scientist whether they had used AI to write a peer-review report, the reviewer admitted using a large language model after running out of time, writing back: "Wow, you guys are good!"
- [7]
Most researchers do not reveal AI help, according to Evanko.
- [8]
GPTZero, a New York City competitor, says it has 99% accuracy and offers "the most precise, reliable AI detection results on the market"; its chief technical officer Alex Cui says five computer-science conferences and three universities have signed up so far, with others piloting.
- [9]
In June, the computer-science conference NeurIPS announced that it rejected 18% of submissions after screening them with Pangram.
- [10]
One in eight biomedical articles last year contained some AI-generated text according to Pangram, in a study reported in January.
- [11]
Users of the preprint server arXiv can check Pangram's verdict on any article there via a mirror site called alphaXiv that has installed the tool.
- [12]
The University of Chicago says it has started using Pangram to vet students' coursework.
- [13]
The tools still sometimes make mistakes, so they can be used only as starting points for investigation.
- [14]
Detector results are less illuminating for AI-edited writing, in which human and AI text intertwine and there is no clear boundary for problematic use.
- [15]
Tim Requarth, who studies science communication at New York University's Langone Health centre, says the detectors are "good for screening out places that are pumping out slop".
- [16]
Pangram co-founder Max Spero has personally called out journalists whom Pangram suggests are using AI and, on one occasion, flagged the Pope's social-media posts as AI-written.
- [17]
In July, Pangram was integrated across the blogging platform Substack, allowing readers to see whether it deems posts to be AI-written.
- [18]
Pangram's advertised 99.98% accuracy implies a 0.02% error rate, or one misclassification for every 5,000 documents screened.
- [19]
GPTZero's advertised 99% accuracy implies a 1% error rate, fifty times Pangram's implied 0.02%.
Sources
1 independent publisher whose own reporting we read for this story.
- nature.comAI-detection tools have made huge leaps forward — how good are they?
2 articles · August 24, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
Entities
- Pangram LabsFollow
- GPTZeroFollow
- American Association for Cancer ResearchFollow
- Daniel EvankoFollow
- Max SperoFollow
- Bradley EmiFollow
- Alex CuiFollow
- Marzena KarpinskaFollow
- Tim RequarthFollow
- Conference on Neural Information Processing SystemsFollow
- alphaXivFollow
- arXivFollow
- SubstackFollow
- University of ChicagoFollow
- Epoch AIFollow
- Simon Fraser UniversityFollow
- New York University Langone HealthFollow
- NatureFollow