Skip to content

Build1 publisher3 min readPublished

Which prompt built the test set decides Pangram's AI-assisted false-positive rate

Pangram's own 4.0 testing put the false-positive rate on AI-assisted documents at 0.01 percent, 4 percent or 7 percent depending on how the test text was produced, and the website shows the smallest of the three.

The Engineer · Build desk

Illustration accompanying Which prompt built the test set decides Pangram's AI-assisted false-positive rate

What happened

  • Pangram's testing of its 4.0 product produced three different false-positive rates on AI-assisted documents, 0.01 percent, 4 percent and 7 percent, depending on which experiment you read.
  • The experiments that returned 4 percent and 7 percent were left off Pangram's website, which advertises that the product detects AI-assisted writing.
  • Until a LessWrong author contacted the company on September 17th, the site put "99.9%+ Accuracy" directly beside "Detects AI Assistance" on the text detection input box.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint The published error rates count a miss only when a whole excerpt is called AI-Generated, so a gate that flags a paragraph or a passage is running outside the regime Pangram measured.
  • decision A buyer now has to pick which of the three test populations resembles the text it will see before quoting any accuracy number in a policy document.
  • exposure Writers who hand a finished draft to a model for a style pass are the ones who land in the 4 to 7 percent bucket, and the post argues those errors cluster by writing style rather than spreading evenly.
  • cost Chunked input raises the error rate on identical text, so the cost of screening short forms such as posts and emails falls on the author being judged.

A false-positive rate on AI-assisted text is a statement about a test corpus, and Pangram built its corpus by prompting Claude to modify human writing. The prompt decides what "AI-Assisted" means. For the experiment that produced 0.01 percent, the instruction to Claude was "Fix spelling, punctuation, and clear grammar errors only" [5]. For the experiment that produced 4 percent, it was "Substantially rewrite for polished academic style while preserving meaning" [6].

The headline figure therefore transfers to your users only if they use models the way the first prompt does. If the people you are screening paste in a finished draft and ask for a style pass, the number from Pangram's own testing that describes them is 4 percent or 7 percent [1].

The version history sharpens this. Pangram released 4.0 on July 29, and the previous version scored 0.2 percent, 15 percent and 22 percent on the same three experiments [7]. The grammar-only rate fell about twentyfold [16]. The two rewrite experiments improved by roughly 3.8 and 3.1 times [17]. "At least an order of magnitude is a big difference! But it may be much worse than that," the post's author wrote [8]. The order of magnitude is in one experiment of the three. Across 4.0's own results the spread from best to worst experiment is a factor of 700 [18]. Pangram published all three; the website shows one [2].

The 4 percent and 7 percent figures counted an excerpt as a false positive only when the entire excerpt was labeled AI-Generated [10]. The author ran about 5,000 words of his own writing through 4.0, and roughly 30 percent of the passages came back AI-Generated, often with high confidence [9]. Many of the longer excerpts were labeled Mixed overall [10]. Under that accounting, Mixed is not a false positive.

His process is on the record. He gives an LLM a prompt like "Please make the wording of the blander parts of this excerpt more poetic and visceral, to match the parts of it that are most in that style. Do not remove core concepts or add additional ones", then spends roughly six more hours per 1,000 words editing and rejects the vast majority of the suggestions [12][11]. At that rate the 5,000-word sample carries about 30 hours of hand editing [19].

Chunking identical text into a few hundred words at a time raises the false-positive rate sharply, and Pangram acknowledges its results are more uncertain for short passages [13][14]. Screening X posts, emails or short answers puts the gate at the length the vendor itself calls least certain [14].

This is one writer's account of Pangram's published experiments, and the post does not include a response from Pangram. He also writes that the errors are likely concentrated in the work of writers whose style confuses the classifier, so a clean result on elite writing does not establish the typical case [15]. He now restricts AI editing to his more private writing [20].

What to watch

  • Whether Pangram publishes the 4% and 7% experiments next to its headline accuracy figure, or restores the adjacency on the input box.
  • Whether Pangram publishes false-positive rates broken out by passage length, given it already concedes short passages are less certain.
  • Whether anyone reproduces the 30% passage-level rate on a corpus that is not self-selected by a single author.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories