Build2 publishers2 min readPublished
Hundreds of contractors read ChatGPT chats behind a filter tuned for identifier fields
OpenAI's outside reviewers read real conversations with account names masked. The 404 Media documents show some review tasks also carry a summary of what ChatGPT has remembered about the user, including location.
The Engineer · Build desk

What happened
- A 404 Media investigation published September 14th found OpenAI using hundreds of outside contractors to read and evaluate real ChatGPT conversations, work the internal documents call Project Lily.
- OpenAI's own technical description of Privacy Filter says the model can miss uncommon identifiers or ambiguous private references and can over-redact or under-redact when context is limited.
- Reviewers score each of four candidate responses from one to seven, where one is unusable and seven is an answer they think would be difficult to improve meaningfully.
- The relevant control, "Improve the model for everyone", is enabled by default and only applies to new conversations, according to the-decoder's account of the investigation.
- Anthropic confirmed to 404 Media that it uses human reviewers for accounts with "Help improve our AI models" enabled, stripping account details such as email addresses first.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure An employee pasting incident detail into the consumer app can put that text in front of an outside reviewer together with a summary of what ChatGPT stored about them, including location, with the account name removed and the situation intact.
- constraint Masking settings leave the problem in place. The review needs the context that makes a chat useful, and per the 404 Media reporting that same accumulated context is what makes an anonymized chat traceable.
- decision Teams that treated consumer ChatGPT as a private scratchpad now have to decide whether temporary chat mode, which OpenAI says is excluded from training, becomes the default for anything work-related.
Privacy Filter works on classes of token. OpenAI says the filter is built to detect and mask names, contact details, addresses, account numbers, passwords and other identifiers before contractors receive the text [4]. Those are fields. A workplace dispute, a medical history or a detailed family problem is prose, and removing an account name does not make such a conversation anonymous, because the content itself points at a person [6].
Some Project Lily tasks arrive with more than the chat. The documents describe a "user memories summary" covering previous interactions, location information and other context ChatGPT has retained about the user [3]. The same accumulated detail that lets the model personalize an answer is what makes a supposedly anonymized conversation easier to connect to a real person [7].
A reviewer summarizes what the user was trying to accomplish, then highlights at least three passages as aligned or misaligned with OpenAI's instructions and explains each call [9]. Recruiting runs through Crossing Hurdles and payment through Mercor, and one North America-based reviewer said they earn more than $50 an hour [17].
The switch is in Settings, under Data Controls, and it stops new conversations from being used for training [14]. Combine that with the default state the-decoder reports, and the backlog is the problem: every conversation typed before the toggle was flipped was already collected under the setting that was on [22]. OpenAI also offers a temporary chat mode which the company says does not use input data for model training [19].
Disclosure sits in the consumer data FAQ, which says authorized personnel and trusted service providers may access user content to improve model performance unless the user opts out, and advises against entering sensitive information [13]. That notice has been online since at least 2023 with only slight changes [18]. One reviewer told 404 Media he did not think users knew humans were reading their chats, and some of the prompts included users explicitly asking ChatGPT to keep the contents private [16]. When 404 Media asked where users are told about human review, OpenAI initially did not respond [24].
The documents leave the scale open: how many users' chats were reviewed, how often the Privacy Filter failed, whether any contractor identified a user [8]; the model being evaluated goes unnamed too [12]. Google notes in Gemini that humans may review saved chats [21].
For a team the re-check is short: whether the toggle is off, when it was turned off, and whether the text people paste would still identify someone after names, addresses and account numbers are masked [14][4].
What to watch
- Whether OpenAI moves the human-review notice out of the FAQ and into the settings screen or onboarding.
- Any published figure for how often the Privacy Filter fails, or how many conversations Project Lily has covered.
- Whether the model being trained under Project Lily is ever named; it goes unidentified in the leaked documents.