Security1 publisher3 min readPublished
Hundreds of OpenAI contractors read whole ChatGPT conversations to score the model's replies
404 Media obtained instruction guides, Slack channels and real prompts from OpenAI's human review program. Reviewers see full chats and a summary of what the user asked for before, with the username stripped.
The Watch · Security desk

What happened
- 404 Media reported that OpenAI is hiring hundreds of contractors to read real ChatGPT prompts, sometimes whole conversations, drawn from a user base the outlet puts at more than 900 million.
- OpenAI said conversations pass through a version of its Privacy Filter model, built to strip personal information, before contractors see them.
- The material 404 Media saw includes instruction guides, Slack channels, the reviewer rating system and real user prompts, which the outlet declined to quote to protect sources.
Compiled by The WatchSomething wrong?How this is made
Why it matters
- exposure Staff who paste client records, contract drafts or incident detail into a personal ChatGPT account are handing that text to a third-party review workflow, because training is on by default on consumer plans.
- constraint An organisation cannot present the pre-review redaction as a technical safeguard to a client or a regulator when the vendor's own documentation lists missed identifiers and under-redaction as expected failure modes.
- decision Acceptable-use policy has to treat consumer chat as disclosure to a human reviewer, not as a scraping-and-training abstraction, and this stream is separate from the safety review OpenAI has already announced.
- precedent With Anthropic confirming human review as well, buyers should assume reviewer access is a category norm and ask for it in writing at contract stage.
A reviewer picks a task from a dashboard and is shown the real prompt of a real ChatGPT user [3]. Sometimes a section above it carries a "user memories summary" describing what that person has previously used the chatbot for, and in some cases where in the world they may live [4]. The username is not there [5]. Anonymization stops at the account label, and the text of the conversation arrives intact [6].
The work runs in three stages: read the prompt, summarize what the reviewer believes the user is asking for, then rate and critique a set of generated responses [7]. One instruction guide 404 Media saw sets the target: "An excellent response should understand the user's intent, provide helpful and accurate assistance, and write in a style that is clear, natural and appropriately warm" [8].
OpenAI's control on the way in is another model. The company told 404 Media that conversations pass through a version of its Privacy Filter, which is designed to detect and remove personal information, before they reach contractors [9]. OpenAI's own page on that model says: "Like all models, Privacy Filter can make mistakes. It can miss uncommon identifiers or ambiguous private references, and it can over- or under-redact entities when context is limited, especially in short sequences" [10]. The vendor describes removal of personal data before human review as best effort [11].
ChatGPT has more than 900 million users, and 404 Media reported that most of them probably do not realize a person may read their conversations [2]. Some of the prompts reviewers see ask ChatGPT to keep the content to itself [12]. "No," a person who works with the prompts said when asked whether users know humans are reading their chats. That person said they did not think users would imagine some contractor somewhere was analyzing the conversations [13].
The reviewers exist because the model has behavioural problems that scraping cannot fix. Internal documents show contractors training ChatGPT not to anthropomorphize itself and to be less sycophantic, a trait that in the 4o model contributed in part to multiple suicides, according to lawsuits cited by 404 Media [14]. This review stream is separate from the chat review OpenAI has announced publicly for users it detects planning to hurt other people [15].
For anyone writing an acceptable-use policy, the relevant line in 404 Media's reporting is that model training is on by default on consumer plans [1]. That account is about consumer tiers; it does not set out how business or enterprise subscriptions are handled. Anthropic confirmed to 404 Media that it also uses human review to improve its models [16], so the same question arises at a different vendor.
Contractors are instructed to escalate tasks they encounter "with potential safety concerns" or personal information [17], and a contractor has read the task before any escalation happens.
What to watch
- Whether OpenAI publishes a user-facing statement that humans may read prompts to improve the model; 404 Media said it put that question to the company.
- Any documentation of whether business and enterprise subscriptions feed the same contractor review pool as consumer plans.
- A data protection complaint that cites the Privacy Filter's own documented under-redaction as the reason de-identification failed.