Inside Project Lily: what it means that humans are reading ChatGPT chats
On 14 September 2026, Joseph Cox reported for 404 Media that OpenAI is paying hundreds of contractors to read real ChatGPT conversations and grade the model's replies. Nothing was hacked. Everything below is what the reporting says, and what it means if your staff are pasting work into a personal ChatGPT tab.

The reporting is based on leaked internal documents, instruction guides, Slack channels, real user prompts, and the rating system the reviewers use. We have not seen that material. Everything attributed below comes from 404 Media's account of it, from OpenAI's own published pages, or from the settings screen on your own account, which you can check yourself in about a minute.
What 404 Media found
Hundreds of contractors read what 404 Media calls "a massive stream of real users' ChatGPT prompts." The prompts can include whole conversations rather than single messages. ChatGPT has more than 900 million users, and, as the piece puts it, most of them probably do not realise a person may read what they typed.
Asked whether ChatGPT users know humans are reading their chats, one of the people doing the work answered plainly: "No. I don't think they would imagine some contractor somewhere is analyzing the conversations."
The work exists to make the model better. Internal documents show contractors training ChatGPT not to anthropomorphise itself and to be less sycophantic, which matters more than it sounds: 404 Media notes that the over sycophantic 4o model has been linked in multiple lawsuits to several people's suicides.
The instruction guide states the target plainly. "An excellent response should understand the user's intent, provide helpful and accurate assistance, and write in a style that is clear, natural and appropriately warm."
How the work actually runs
Reviewers log into a dashboard and pick a task. The task opens with a real user's prompt, and the job runs in three stages: read the prompt, summarise what the user was trying to do, then rate and critique a set of responses ChatGPT generated for it.
The summary step is exactly what it sounds like. One example in the guide reads: "The user is asking for help on revising a work Slack message. They want it to sound collaborative and invite input from tagged people."
Then the reviewer reads four candidate responses and highlights at least three passages as aligned or misaligned with the model being trained, each with a written reason. The guide's worked example marks a list decorated with the ✅ emoji as misaligned, reason given as unnecessary use of emojis. Another document says an assistant register and "emoji misuse" pull scores down when they hurt the user's experience, and that context decides: a tree emoji is fine for Arbor Day, skull emojis in a conversation about death are not.
The same material tells the model to avoid claimed experience, no "As a chef, I like to" and no "I know what that's like," while ordinary first person such as "I'll take a look" is allowed.
Finally the reviewer scores each response from one to seven. One is described as unacceptable and unusable. Seven is described as a response that would be hard to meaningfully improve. A document marked "Confidential & Proprietary" tells reviewers the model should "generally match the user's tone, but slightly less intensely," and should "remain natural, restrained, and professional without implying that it is human or experiencing emotions."
Fact checking is not the job. An FAQ tells reviewers OpenAI does not expect them to verify claims with outside searches, though they are asked to flag correctness problems they happen to notice and to penalise missing sources on high stakes medical, legal and financial answers.
Put the whole thing together and the shape is clear: a person reads what you wrote, writes down what they think you wanted, and grades how the machine spoke to you. Tone is the product. Your situation is the raw material.
Your name is removed. That is not the same as anonymous.
The dashboard does not show the ChatGPT username. OpenAI told 404 Media it runs conversations through a version of its Privacy Filter model first, designed to detect and remove personal information.
OpenAI's own page describing that model is unusually direct about the limits, and it is worth reading slowly rather than skimming:
"Like all models, Privacy Filter can make mistakes. It can miss uncommon identifiers or ambiguous private references, and it can over or under redact entities when context is limited, especially in short sequences."
Uncommon is the word to sit with. The things that are unusual about you are the things most likely to survive a general purpose redaction model, and unusual is precisely what identifies a person.
What the filter is not trying to remove is the situation itself, because the situation is the thing being graded. And above the prompt there is sometimes a "user memories summary," an overview of what that user has previously used the chatbot for, which 404 Media says can include where in the world the person may live and other personal context about them.
A role, a region, an employer and one specific problem is not a description of nobody. Reviewers are told to escalate tasks with potential safety concerns or personal information, which is a sensible instruction and also an admission that both arrive.
One detail in the reporting is harder to shake than any of the policy language. Some of the prompts reviewers see are ones where the user asks ChatGPT to keep the contents to themselves.
The setting is on unless somebody bought you the other one
OpenAI told 404 Media that chats are not used to improve its models if the user turns off "improve the model for everyone."
That setting is on by default for Free, Plus and Pro. It is off by default for Enterprise, Business and Edu. Turning it off applies to new conversations and does not appear to work retroactively.
Read those three sentences together, because they are the whole security argument in this story. The accounts a company signs a contract for are excluded from this pipeline before anybody touches anything. The accounts people open for themselves are included by the same default. Whatever your data processing agreement, your zero retention terms, your subprocessor list and your SOC 2 report cover, they cover the tier you bought. Your employees are pasting production logs, customer tickets and half a contract into the other one, from a personal account, in a browser you do not manage, because it is the fastest way to finish.
404 Media asked OpenAI whether it had ever explicitly told users that humans may review their prompts. OpenAI did not answer the question. After the company was contacted for comment it updated its help page about the setting with more detail on opting out, and 404 Media reports it still does not acknowledge that humans may read users' prompts. After publication, OpenAI pointed to a section of its site saying humans may review content to improve model performance.
On deletion, OpenAI says conversations leave its systems within 30 days, unless the content has already been stripped of identifiers and, in its words, "disassociated from your account when you allow us to use your Content to improve our models."
Here is how to switch the setting off. It takes four taps, and the row behind the dialog has to read Off before anything has changed.

Who is actually reading it
This is the part most coverage skipped, and it is the part that decides how much the word "contractor" should reassure you.
The worker 404 Media spoke to lives in North America and is paid more than fifty dollars an hour. They found the work through a recruitment firm called Crossing Hurdles, whose website says it "connects skilled professionals with AI training, evaluation, research, and contributor opportunities across the global AI economy" and, in a line that reads differently after this story, "Human intelligence powers AI progress." 404 Media notes multiple people on Reddit reporting unsolicited recruitment emails from the firm, some asking each other whether it was a scam.
Crossing Hurdles refers people on to Mercor, an AI training company, which is who actually pays the people reading the prompts. 404 Media also reports that Meta stopped working with Mercor in April, after Mercor suffered a massive data breach.
That is the sentence to hold on to. The chain that ends with your conversation on somebody's screen runs through at least two companies most users have never heard of, and one of them has already been breached.
The worker described the job as sometimes "kind of amusing" and on the whole "very rote," with guidelines that change often and can feel self contradictory.
None of this is new as an industry practice. A TIME investigation found OpenAI hired Kenyan workers to label text in order to make its platform less toxic. Google's Gemini carries a disclaimer saying humans review some saved chats to improve Google AI. Anthropic told 404 Media it also uses human review to improve its models, for users who have turned on the "Help improve our AI models" setting, and says it removes account identifiers such as email addresses before review.
Michal Luria of the Center for Democracy and Technology put the mismatch precisely: chatbot interfaces "automatically create a false sense of intimacy and privacy" in what feel like private exchanges, "when in reality there may be human reviewers reading on the other end." She notes this is quite distinct from social media, where posting already carries an expectation of moderation.
Sarah T. Roberts of UCLA, who wrote a book on content moderation, compared it to the Wizard of Oz, "where the protagonists discover that the magical kingdom is really a man behind a curtain pulling levers." On the pay, she was blunt that the rate is what it is "for now," and that the human work these products depend on is the work companies pay least for and value least.
What this means for a security team
Nothing in the story is a breach, and that is the problem. There is no incident number, no disclosure clock and no post mortem, because from the provider's side nothing went wrong. The data was handed over, under a setting that was already on, into a documented process, and read by somebody under contract.
Every control most companies bought sits upstream or downstream of the only moment that mattered. Your network tooling sees a visit to a chat site, not the four thousand characters of customer data inside the request body. Your enterprise agreement covers an account your employee is not signed into. And both of the switches a user can reach, opting out and deleting, attach after the conversation has already been sent.
So look at what is actually left, because it is not nothing and it is not a product pitch yet.
Every remaining control in this story has the same shape. Turning the setting off, using Temporary Chat, clearing memory, describing the problem instead of pasting the log: each one is a decision a person has to make correctly, in the moment, on whichever device they happen to be holding. Train people on all of it and you have improved the odds. You have not changed the shape.
And the moment these controls are needed is the moment they are least likely to happen. Nobody pastes an entire production log because they think it is a good idea. They do it at eleven at night because describing the problem properly would take fifteen minutes and pasting takes four seconds. Awareness training is asking the tired version of somebody to outperform the rested version, which is not a control, it is a hope with a slide deck.
SolonGate Shadow AI exists because that gap is structural. It reads the content in the browser before it is sent, on every tab, for everybody, and it does not ask anyone to remember anything. It blocks, or it redacts, or it lets the paste through and writes down what left, which is the part most organisations discover they never had: an answer to the question of what has already gone.
It is deliberately not a proxy and not a domain blocklist. Both lose the moment somebody opens a tab nobody anticipated, and banning the tool outright just moves the paste to a phone, where nothing is watching at all.

What we can measure about it
A policy cannot be measured. A scanner can, and that is the only reason to prefer one. So in September 2026 we published the benchmark behind ours, with the methodology and corpora in the repository so anyone can rerun it and argue with the result.
Against 1,890 real pastes carrying 4,930 marked values, the scanner finds 94.7 percent. Against a clean corpus of 3,642 pastes and roughly 1.1 million tokens of real open source code, it raises 0.25 false alarms per 10,000 words. The second number is the one people skip: a scanner that interrupts a developer forty times a day gets switched off in a week, and a scanner that is switched off detects nothing.
We published the rows where we lose, and we are not hiding them here either. Our passport detector produced 29 false findings out of 50 in an early run. The benchmark found five bugs in our own scanner, and fixing them moved detection from 76.4 percent to 94.7 percent.
What none of it fixes
Shadow AI cannot reach into somebody else's review queue. If a conversation has already been sent and flagged, it has been sent and flagged. It will not stop a person retyping a customer's situation from memory in their own words, because there is no marked value in that sentence to find. It will not turn 94.7 percent into 100 percent.
And for you personally, the honest summary is shorter than most advice on this subject: turn the setting off, clear your memory, use Temporary Chat for anything you would not want read, describe the problem instead of pasting the artifact, and use the account your employer pays for when the question is a work question. Each of those helps. None of them is anonymity, because the one thing no filter removes is the reason you opened the box.
Sources. Joseph Cox, "Inside 'Project Lily': The Humans Reading Your ChatGPT Chats," 404 Media, 14 September 2026, which is the original reporting and the source of every internal document, quote and figure attributed above. OpenAI's own Privacy Filter, data controls and privacy pages for the quoted model limitations, the default settings by tier and the 30 day deletion language. Euronews Türkçe, 17 September 2026. Additional coverage: Gadget Review, IBTimes, securityonline.info. Prior reporting referenced by 404 Media: TIME on OpenAI's Kenyan data labellers. If you want the story itself rather than our reading of it, read the 404 Media piece.