Science1 distinct publisher3 min readPublished
Cancer-journal editors are already catching undisclosed LLM use in peer review with commercial software. The unresolved question is what happens to a flagged researcher who says no.
The Scientist · Science desk
Compiled by The ScientistSomething wrong?How this is made
The arithmetic in the marketing is worth doing before anyone drafts a policy. Pangram's website advertises 99.98% accuracy [4], which permits one error for every 5,000 documents screened [1]. GPTZero advertises 99% [5], which permits one in a hundred, fifty times as many [2]. Two products in the same market are not separated by a factor of fifty in quality. They are separated by what was measured and on whose text. Both figures are the firms' own, and Nature's reporting puts the independent view more narrowly: the tools call human-only text human almost all the time [10], and a flag is a starting point for investigation rather than a finding [11].
The AACR case shows what resolution looks like in practice. Daniel Evanko asked, and the reviewer confessed to using a model after running out of time [2]. He did not expect that, because most researchers do not disclose AI help [3]. Take the confession away and the editor is holding a probability and a denial, which is the part of the workflow nobody has designed. The person being scored cannot inspect the system doing the scoring, and proving you wrote your own sentences is not a task with a defined end.
The disputes will not mostly involve wholly synthetic text. Pangram's reading, in a study reported in January, put some AI-generated text in one in eight biomedical articles last year [7], and the tools are weakest precisely where that volume sits: writing in which human and machine prose intertwine, with no clean boundary for what counts as problematic use [12]. Tim Requarth of NYU Langone says the detectors are good for screening out places pumping out slop [13], and that case is easy. Deciding whether a model that tightened a reviewer's grammar breached a disclosure rule is a definitional question, and a score does not answer it.
Opting out is not on offer either. Pangram verdicts sit on arXiv papers through the alphaXiv mirror [8], on Substack posts since July [18], and on coursework at the University of Chicago [9]. Its co-founder Max Spero has publicly flagged journalists and, on one occasion, the Pope's social-media posts [17]. Complaints will therefore reach research-integrity offices from third parties whether or not the office screens anything itself. The accused will have a ready defence in the record: before this generation of tools, detection software was notoriously unreliable, in the account of Simon Fraser University computer scientist Marzena Karpinska [14], because the statistical approaches flagged human-authored text as AI too often [15]. An office that wants its flags to carry weight has to publish its threshold and its appeals route before it uses one in anger.
Ranked by verification strength, evidence, and original report placement.
Daniel Evanko, director of journal operations at the American Association for Cancer Research (AACR), and the AACR deploy the commercial AI-detection tool Pangram because of concerns over the number of peer-review reports submitted to their journals that appear to use AI without disclosure, contrary to the publisher's policy.
Pangram Labs, the New York City start-up behind Pangram, states on its website: "Detect AI-generated content with 99.98% accuracy."
Independent analysts say the detectors do work in the sense that they correctly flag solely human-written content as human almost all the time, although no tool can be perfect.
Until Pangram and others came along, AI-detection software was notoriously unreliable, says Marzena Karpinska, a computer scientist at Simon Fraser University who has conducted independent analyses of the tools.
Earlier statistical approaches, based on perplexity, burstiness and stylistic quirks such as over-use of the 'It's not X, it's Y' construction, too often falsely flagged human-authored text as AI-generated.
When Evanko asked a scientist whether they had used AI to write a peer-review report, the reviewer admitted using a large language model after running out of time, writing back: "Wow, you guys are good!"
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed single-publisher reporting with named independent testers, but headline accuracy rests on vendor figures
The account is specific and attributed — a named journal-operations director, a verbatim reviewer admission, named deployments, and independent researchers (Karpinska; an Epoch AI July report of zero false positives) corroborating low false-positive behaviour. Against that: the cluster is a single article ingested twice, the 0.02% false-positive figure originates in a non-peer-reviewed vendor study, the supplied body is truncated mid-sentence, and the limits of the tools are asserted qualitatively rather than quantified.
Multiple named production deployments across publishing, conferences, a university and consumer platforms
Adoption is concrete rather than announced-intent: AACR journals screening peer review, NeurIPS rejecting 18% of submissions after screening, University of Chicago coursework vetting, alphaXiv exposing verdicts on arXiv, and a platform-wide Substack integration. GPTZero adds five conferences and three universities. The absolute counts remain small and no seat-count, volume or revenue figures are given, which is why this is not higher.
Marketing precision outruns the governance and mixed-authorship caveats the same reporting concedes
Vendor language ('99.98% accuracy', 'most precise, reliable AI detection results on the market') is meaningfully stronger than what the reporting substantiates: the tools still err, are usable only as starting points, and are least informative exactly where real disputes arise — AI-edited text with no agreed problematic-use boundary. Independent corroboration of low false positives keeps the gap modest rather than large, and the deployment evidence is real, so this is overstatement at the margin rather than vapour.
Commercial detector vendors, enforcement-seeking publishers, and a publisher-owned outlet all have stakes
Two competing New York start-ups market on precision and stand to gain from norms that make undisclosed AI use detectable and sanctionable; one co-founder publicly names individuals his tool flags and explicitly frames the moment as norm-setting. Buyers — journals, conferences, universities — need defensible screening to enforce disclosure policies. The reporting outlet is itself a scientific publisher covering tooling sold to scientific publishers, and prevalence statistics that grow the market come from a vendor's own analysis.
Well-sourced but single-publisher and partly vendor-dependent
Confidence is limited by cluster structure: one publisher, two identical ingests, and a truncated body that cuts off mid-sentence in the Epoch AI passage. Within those limits the reporting is specific, names its sources, quantifies deployments, and states its own caveats, and independent researchers are cited alongside vendors — so the core factual spine (deployments happened; tools are materially better than the prior generation; residual error and mixed-authorship ambiguity remain) is reasonably firm, while precise accuracy figures and customer counts remain unverified.
science
Pasqal's prompt-to-circuit agent still needs a physicist in the loop1 distinct publisher
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
science
Rydberg chain spectra match Ising CFT predictions, turning a simulator into an instrument1 distinct publisher
science
GJ 523b gives 'Mega-Earth' a number: 23 Earth masses inside 2.5 Earth radii1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
2 articles · August 24, 2026