Skip to content

Security1 publisher2 min readPublished

Reviewers policing OpenAI's training contractors are barred from using AI detection tools

404 Media found contractors who read real ChatGPT conversations have been fired for handing in model-generated work, and the internal documents show the sanctioned check on them is a human judging the overall pattern.

The Watch · Security desk

What happened

  • 404 Media, working from internal documents and interviews with three contractors, found that people hired to improve OpenAI's models have been fired for using AI to produce their training work.
  • An internal document tells the contractors who review other contractors' work not to use GPTZero or any other AI detection tool because they are not reliable, and not to use AI themselves, including Grammarly and AI translation.
  • Project Lily, which 404 Media reported the previous week, has hundreds of contractors reading real ChatGPT users' prompts and conversations that can include personal information.
  • Mercor, which employs two of the contractors who spoke to 404 Media, said it removes an expert from a project as soon as it confirms that person used AI to complete a task.

Compiled by The WatchSomething wrong?How this is made

Why it matters

  • constraint Nobody at the vendor or the client can say where an individual annotation came from. The only permitted test is a reviewer's judgement about overall pattern, and it runs after the work is handed in.
  • exposure Read access to real user conversations sits with an outsourced workforce that can get through a task faster by using the AI its contract forbids.
  • decision Anyone buying human-in-the-loop review has to decide what a human-origin attestation is worth when the supplier's own guidance calls the available detectors unreliable.
  • contradiction Mercor says it invests heavily in tools and systems to detect misuse. The internal document tells its reviewers that detection tools are unreliable and bars them from using them.

The enforcement loop is built to stay opaque to the person being judged. "Do not tell evaluators why you suspect AI," the document for reviewers says. "It is easier for them to hide if they know what you look for. Judge the overall pattern, not one clue." [8]

The pattern is repetitive words, AI-style punctuation including heavy use of the em dash, and work that comes back very quickly [10]. In Slack channels contractors post samples and ask each other whether something is AI. "Usually the answer is yes," one contractor said [13].

One of the internal documents puts more than ten thousand contractors on these projects [4]. That is at least ten times the headcount 404 Media described on Project Lily [1].

One contractor said they see people using AI "all the time and people are let go for it all the time, it's pretty much the one thing that will get you kicked off ASAP" [11], and that "in a group of thousands there are tons that have been caught" [12]. Two of the three contractors said people have been fired or offboarded [9]. The contractors spoke anonymously because they were not permitted to talk to the press [23]. 404 Media did not report how many were fired.

Mercor said its contracts "strictly prohibit the use of LLMs to complete projects and we enforce that" [19], and that "When we confirm an expert has used AI to complete a task, we immediately remove them from the project" [20]. Both of those statements cover cases Mercor confirmed; certifying that a delivered batch is human-written is a different claim, and the only sanctioned test for it is a reviewer's judgement about overall pattern, applied after the work is submitted [10].

One contractor said they used AI while helping train OpenAI's models and shared what they presented as their termination letter, which said their employer had identified issues with the "authenticity" of their work [14]. "I just needed a little boost and turned to AI to help me which eventually led to my downfall," the contractor told 404 Media [15].

404 Media frames the quality risk as model collapse, where models trained further on AI-generated text get worse [24]. 404 Media also spoke to a fourth contractor, who has worked on training models for various AI companies and said they sometimes purposefully chose the worst responses [22].

What to watch

  • Whether OpenAI or Mercor publishes a rate: how much delivered annotation has been discarded as model-generated, and over what period.
  • Whether either company states a re-check method for work already in the training set, when reviewers are barred from using detection tools.
  • Whether the contractor access to real conversations containing personal information draws a data-processor question from any customer or regulator.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories