Skip to content

Product1 publisher3 min readPublished

The founder of a deepfake-detection firm verifies his wife's calls with a safeword

Two detection chief executives describe their own products as probabilities rather than verdicts. That is the constraint to design around before a score is allowed to fail a student or close an account.

The Product Desk · Product desk

Photograph accompanying The founder of a deepfake-detection firm verifies his wife's calls with a safeword
Photo: getrealsecurity.com

What happened

  • Hany Farid, a Dartmouth computer science professor, founded the deepfake-detection company GetReal on the same day ChatGPT was publicly released in November 2022.
  • Farid and his wife now use an agreed safeword on remote calls, after a scammer cloned his voice to try to extract sensitive information from a lawyer he was advising.
  • Pangram, started in 2023 by two Stanford graduates, claims to flag AI-generated text with 99.98% accuracy, and an independent University of Chicago comparison rated it the most effective tool tested.
  • Pangram chief executive Max Spero told Gizmodo he does not think detection will ever reach 100% accuracy, and that the most his company can do is accumulate more certainty from more data.
  • Jon Gillham of Originality.ai likened text detection to a weather forecast, calling it highly accurate but not perfect, and said he would bet on that staying true.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • cost The error rate is spread unevenly: the individual who is wrongly flagged carries the accusation and the burden of disproving it, while the savings from automatic screening accrue to whoever installed the detector.
  • constraint Because a tell can disappear when a model vendor changes a default, any threshold tuned this quarter has an unknown expiry, which makes a detector a standing revalidation cost rather than a one-time purchase.
  • decision Anyone attaching a penalty to a detector output now has to choose whether it is reversible, since the vendor's own on-record position on certainty is available to the first person who contests it.
  • capability A pre-agreed word between two people gives them a verification step that does not weaken as generative models improve, which is something no confidence score offers.

Turn Pangram's headline accuracy around and it reads differently: a 0.02% error rate is one wrong call in every 5,000 documents [1]. A writing programme pushing 20,000 submissions through a term should expect roughly four people on the wrong side of it [2]. Gizmodo's account does not say how that 0.02% divides between wrongly accusing a human and quietly clearing a bot [16], and those two failures have different owners. One of them produces a hearing; the other produces nothing, which is why it never reaches anyone's complaint queue.

The number also has no stated shelf life. Em dashes were the tell everyone had learned, and then Sam Altman posted in November 2025 that ChatGPT would finally obey a custom instruction to stop using them [15]. A habit that heuristics were keyed to thinned out on someone else's release schedule. Detectors sit downstream of that, and downstream of humanizer sites whose stated purpose is reworking AI text until it reads as human [14].

The claim that generation is outrunning detection is worth locating precisely. Hany Farid told Gizmodo that generative AI "is growing up very, very quickly" and "is being weaponized" [5], and that he now mistrusts much of what he sees on a screen [6], which is unease rather than a measurement. The measured version comes from the text side, where detection is hardest because generated text follows much simpler statistical patterns than photorealistic images or video [7]. Both of the hedges on record came from vendors describing their own products, which is the part a buyer should read twice.

For a team wiring a score into a product, the benchmark ranking matters less than what the score is permitted to trigger. Two axes sort it. Does the action reverse, and does the flagged person see the evidence and get to answer it? Reversible and contestable is where a probability belongs: queue the post for human review, ask the candidate to walk through their draft, put the clip in front of someone who can look at the artefacts. In the irreversible corner, meaning a failing grade, a closed account, a denied claim, a single model's score needs corroboration that did not come from a model, because the vendor's own public position on certainty is the first thing an appeal will quote [10].

The forcing function is a two-name test: the person who gets wrongly flagged in your system, and the person who answers their email. When the second name does not exist, the score is doing enforcement work nobody has agreed to own. Farid's fallback, one agreed word between two people [4], keeps working after the next model ships, because it never depended on recognising a model's output in the first place.

What to watch

  • Whether Pangram or Originality.ai publish false-positive rates separately from headline accuracy, and name the corpus.
  • A rerun of the University of Chicago comparison against current models and current humanizer tools, which would show how fast the ranking decays.
  • The first appeal in which an institution's irreversible penalty is contested using the detection vendor's own statement that it will never reach 100% accuracy.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories