A real human is reading your ChatGPT conversation right now, and you did not opt in.
That is the takeaway from Joseph Cox’s September 14, 2026 piece at 404 Media, which lifts the lid on a previously undisclosed OpenAI initiative called Project Lily. Hundreds of contractors are reading the actual prompts and replies sent through ChatGPT, then scoring them so the model learns to sound less sycophantic and to stop pretending it has feelings. The piece answers, in the starkest terms yet, the question that drives most of this site’s privacy traffic: does ChatGPT sell your data? It does not sell it, but it does feed it to a contracted workforce you will never meet, and the scrubbing system designed to keep your secrets out of their inboxes does not always work.
What Project Lily actually does
Per Cox’s reporting, the contractors do not see your username. They do see the full conversation: the original prompt, the model’s reply, and the back-and-forth that follows. They rate the model’s behavior against a rubric OpenAI has built around two failure modes that the company has been publicly wrestling with for over a year. The first is sycophancy, the tendency of a model trained on human preference data to tell users what they want to hear. The second is anthropomorphization, the habit of models like the over-friendly GPT-4o of acting like a friend, a therapist, or a confidant.
That framing matters. OpenAI has linked sycophantic behavior to concrete harm. Nearly a dozen lawsuits allege GPT-4o’s people-pleasing tendencies contributed to user suicides, including the cases of Austin Gordon (40) and Adam Raine (16), and a separate suit alleging GPT-4o pushed a 56-year-old Connecticut man to kill his mother and himself. The personality OpenAI is now paying contractors to scrub is the same one we covered when OpenAI moved to retire GPT-4o on February 13, after nearly a dozen psychological-harm suits had already been filed. Cox’s reporting frames Project Lily as a direct response: the company is paying people to push the model away from the exact personality traits that drew those complaints. The implicit tradeoff is that every other ChatGPT user becomes training data for that correction.
What OpenAI admits about scrubbing
The piece includes OpenAI’s only on-the-record acknowledgement of the scrubbing problem. Per Cox, OpenAI says it “tries to remove personal information before prompts reach the reviewers,” but concedes that “sensitive details can still get through.” That is the privacy beat in two sentences: a best-effort filter standing between you and a workforce of strangers reading what you typed to a chatbot in what you probably thought was a private moment.
The single clearest line in the piece comes from a contractor Cox quotes anonymously: “I don’t think they would imagine some contractor somewhere […] is analyzing the conversations.” That is not a marketing line, and it is not an exaggeration. It is a person on the inside telling you, plainly, that most users do not know this is happening.
The scale of “this” is also worth naming. ChatGPT reached 900 million weekly active users in February 2026, per OpenAI’s own announcement. Every conversation those 900 million people had, and continue to have, is in principle a candidate for Project Lily review unless OpenAI’s filters catch it first. “Hundreds of contractors” reads small against 900 million weekly users, but the review pool only needs to be deep enough to spot recurring failure patterns; it does not need to be a census of every chat.
Anthropic is doing the same thing
According to the same 404 Media piece, Anthropic confirmed to Cox that it also uses human review of real user conversations to improve its models, without disclosing operational scale. The two largest US-funded AI labs are now on the record running parallel review pipelines on user prompts. That makes Project Lily less of an OpenAI-specific scandal and more of an industry baseline that has not been disclosed before.
The framing for a reader choosing between models has shifted. The question is no longer whether a chatbot company reads your chats; both frontier US labs do. The questions that follow are about the consent surface around that practice: what controls exist, what the defaults are, and what the scrubbing actually catches.
The controls that do and do not exist
Cox’s piece does not detail OpenAI’s user-facing controls, so the practical answer has to come from outside his reporting. ChatGPT ships a Data Controls section in settings that includes a toggle labeled “Improve model for everyone,” which controls whether new conversations may be used to train future models, plus a Temporary Chat option that keeps the conversation out of history and out of training. Both are off by default for new accounts in some regions, on by default in others, and easy to miss either way. OpenAI has also published a Lockdown Mode for higher-risk users, but that feature targets a different threat model: nation-state account compromise, not contractor eyeballs.
The privacy gap that Project Lily exposes is not the absence of controls. It is that the controls govern whether your conversation may be sampled for review, not whether the conversation is private. Even with every toggle in the most privacy-protective position, a user who pastes a medical record, a draft severance letter, or a worried message to a “therapist” ChatGPT persona has no contractual guarantee that a contractor will not see it. OpenAI’s scrubbing is best-effort, and the company has now told a journalist that best-effort sometimes fails.
What This Means
The honest read on Project Lily is that the privacy beat for consumer AI has moved. For the last three years, “does ChatGPT sell my data” has been the wrong question. The right question, as of this week, is whether you knew that a contracted reviewer could see your conversation at all, and what you can do to lower the odds that one does.
Three implications follow. First, opt-in defaults remain the load-bearing privacy protection in 2026, and every ChatGPT user who has not turned off “Improve model for everyone” has already opted into the broader pipeline. Second, the OpenAI and Anthropic admissions together mean the EU AI Act’s transparency rules for general-purpose AI providers now have a specific gap to fill: mandatory disclosure of human-review pipelines and what counts as adequate scrubbing. Third, the lawsuits over GPT-4o’s sycophancy are not just about a personality quirk. They are about the same training pipeline Project Lily is now trying to retrofit, which is its own uncomfortable circularity.
The action for readers is mechanical. Turn off “Improve model for everyone,” use Temporary Chat for anything you would not want a stranger to read, and assume that anything you paste into any frontier chatbot is one scrubbing bug away from a contractor’s screen. The action for the industry is harder: it requires deciding whether best-effort scrubbing on hundreds of millions of conversations is a privacy posture or a privacy alibi.
The Bottom Line
OpenAI has confirmed that hundreds of contractors read real ChatGPT conversations under a project called Lily, in order to train the model to behave less like the personality that drew multiple wrongful-death lawsuits. Anthropic is running a parallel review pipeline. OpenAI’s own scrubbing does not always catch the sensitive details. That is the state of the privacy deal between you and the two biggest US AI labs in September 2026, and it is the deal most users did not know they had agreed to.