AI & Privacy

OpenAI Sued Over 'Project Lily' Human Review of ChatGPT

Aditya Kumar JhaAditya Kumar JhaLinkedIn·September 28, 2026·11 min read

A new lawsuit alleges OpenAI let contractors read ChatGPT chats undisclosed. What 'Project Lily' means for AI privacy claims, MemX included.

A new federal lawsuit alleges that OpenAI routed real ChatGPT conversations, including chats about health conditions, debt, and legal trouble, to outside contractors who read, summarized, and scored them, without disclosing that a human being might be the one reading. The case, Vredenburgh v. OpenAI OpCo, LLC, was filed September 16, 2026 in the U.S. District Court for the Northern District of California and served on OpenAI on September 21. It centers on an internal program reporters have named 'Project Lily.'

This did not start as a lawsuit. 404 Media broke the story on September 14, 2026, after reviewing internal OpenAI materials, including reviewer instructions, Slack messages, a scoring rubric, and real user submissions. Contractors are recruited through a staffing firm called Crossing Hurdles, paid through the AI-training company Mercor, and reportedly earn more than fifty dollars an hour to read anonymized excerpts of live conversations, summarize what the user was trying to accomplish, and rate multiple chatbot responses on a scale of one to seven. Internal accounts described in the reporting suggest even people inside OpenAI didn't expect users to realize a contractor might be analyzing their conversations.

Two California ChatGPT users brought the suit on behalf of a proposed nationwide class, case number 3:26-cv-10527, assigned to Magistrate Judge Alex G. Tse. OpenAI's response is due October 13, 2026, so none of this has been decided by a court, and OpenAI has not been found liable for anything. This is a different dispute from the New York Times copyright litigation that forced OpenAI to preserve, and later produce, 20 million de-identified chat logs, a separate case with its own timeline. Project Lily has nothing to do with litigation-driven data retention. It is about whether OpenAI told users, up front, that people outside the company would read their actual conversations to grade the model's answers.

Keep two things separate here. That contractors exist, that they are paid to read real conversations, and that the practice happened is independently documented reporting, and OpenAI has not disputed it. Whether that specific practice was inadequately disclosed, and whether it violated California privacy and consumer-protection law, is a legal claim a court has not yet ruled on. A complaint states one side's case. OpenAI is entitled to respond, and courts routinely narrow or dismiss claims long before trial.

What made contractors nervous enough to talk was not the existence of review itself; it was what showed up in the queue. Reporting on the internal materials describes prompts touching therapy-style conversations, professional and legal questions, and in some cases a 'user memories summary' that surfaced details from earlier chats, including location information, none of which the person reviewing that single conversation was supposed to see in isolation. Some of the reviewed material came from users who had explicitly asked ChatGPT, in the conversation itself, to keep something private. Anonymization removed names and account identifiers. It did not remove the substance of what people had written.

What the lawsuit alleges, and what the reporting independently confirms

  • According to the complaint, OpenAI's Privacy Policy lists eleven categories of outside companies that receive user data, and none of them is described as a data-labeling, annotation, or human-evaluation vendor.
  • Reviewers read full conversations, not isolated fragments, and 404 Media's review of internal materials found that OpenAI's automated privacy filter, meant to strip identifying details before a human sees the text, does not catch everything.
  • Reviewers work the queue under job titles like 'AI data reviewer' and 'chatbot evaluator,' picking tasks from a dashboard of real prompts, reading multiple ChatGPT responses per task, and writing a rationale to go with each one-to-seven score.
  • The complaint asks a court to require opt-in consent before any conversation goes to an outside reviewer, default the 'Improve the model for everyone' setting to off, add an in-chat warning, and order deletion of reviewer work product tied to reviewed conversations.

Other AI chatbots use human review too, this lawsuit is about disclosure

OpenAI is not the only company that has people look at real chatbot conversations. Reinforcement learning from human feedback, the technique behind why modern chatbots sound conversational instead of robotic, generally needs human judgment somewhere in the loop, and 404 Media's reporting notes that Anthropic has confirmed it also uses human review to improve its models. The legal fight here isn't over whether any human should ever see a chat. It's over whether this specific pipeline, run through an outside staffing firm at this scale, was clearly disclosed before it started.

What the complaint is actually suing over, legally

A complaint has to hang its facts on a legal theory, and reporting on the filing describes eight stacked California claims, including alleged violations of the state's Unfair Competition Law and Consumer Privacy Act, plus common-law and state-constitutional invasion-of-privacy theories, on top of the core argument that OpenAI's disclosures fell short of what state consumer-protection law requires before a company routes someone's private writing to a third party. None of that has been tested in court yet. No class has been certified, no discovery has forced OpenAI to produce its own account of what the Privacy Policy was meant to cover, and a motion to dismiss, if OpenAI files one, could narrow the case considerably before it ever reaches a jury.

What people assume versus what's actually disclosed

QuestionWhat people assumeWhat OpenAI disclosesWhat the lawsuit alleges
Who reads a submitted chatOnly the AI model, nobody elseAuthorized personnel and 'trusted service providers' may access content to improve model performance, per OpenAI's help centerOutside contractors, hired through a staffing firm and paid by the hour, read full conversations as routine production work
Where this is written downSomewhere in the privacy policy, in plain languageA help center article on model-performance data use, separate from the Privacy Policy's list of data recipientsThe Privacy Policy's list of third-party recipients names no data-labeling or human-evaluation vendor category, the complaint says
How to turn it offThere is probably a simple toggle, and it defaults to offSettings, then Data Controls, then 'Improve the model for everyone', or 'do not train on my content' in the Privacy PortalThe toggle exists, but it defaults to on for Free, Plus, and Pro accounts; the lawsuit wants that default flipped

OpenAI is not silent on this. Its help center carries an article on how conversation data is used to improve model performance, stating that a limited number of authorized personnel, and 'trusted service providers' bound by confidentiality obligations, may access content for reasons including model improvement, unless a user opts out. That page exists and predates the lawsuit. The dispute is narrower than 'OpenAI never said humans read chats.' It's about whether a help center article, describing access in the abstract, is the same as disclosing that a named contractor workforce, hired through a staffing agency and paid by the hour, would read full conversations as a matter of routine work. The complaint says no, and points to the Privacy Policy's own list of third-party recipients, which it says does not include that category.

Pro Tip

Check your own account: in ChatGPT, open Settings, then Data Controls, and look at whether 'Improve the model for everyone' is on. It defaults to on for Free, Plus, and Pro accounts. Turning it off, or choosing 'do not train on my content' in OpenAI's Privacy Portal, stops new conversations from entering this pipeline going forward. It does not retroactively pull back anything already reviewed.

What 'private by architecture' actually means

Every AI company's privacy language says some version of 'we protect your data.' That sentence does no real work until you ask a narrower question: is this privacy a promise written into a policy that can be reinterpreted, updated, or simply not cover some new internal program, or is it a property of how the system is actually built? 'Private by architecture' is a claim of the second kind. It describes technical facts about a system, not intentions about how a company plans to behave.

  • Encryption at rest: stored content is unreadable without the right keys, not just protected while moving between a device and a server.
  • Customer-managed encryption keys (CMEK): the keys that decrypt data are managed per customer, rather than pooled in a way that lets one internal key holder access everyone's content at once.
  • Per-user isolation: one account's data is walled off from another's at the storage and processing layer, not only by an access-control policy someone could misconfigure.
  • On-device processing where applicable: some analysis happens locally, so raw content doesn't have to leave the device to become searchable.

None of that is end-to-end encryption, and none of it is a zero-knowledge system. Those are different, stronger claims, meaning not even the company running the service can read the content under any circumstance. 'Private by architecture' is a narrower, more honest claim: no undisclosed pipeline routes content to outside contractors as routine business, and the technical safeguards around storage and access are structural rather than promised in a policy document that can quietly fall out of date.

This distinction is exactly what the Project Lily lawsuit turns on. Nobody is arguing that OpenAI's servers were breached or that encryption failed. The allegation is about a business process built on top of properly stored data: real conversations, routed on purpose to people outside the company, for a use the complaint says was never clearly named. A policy sentence about protecting data says nothing about which internal or third-party workflows that data is allowed to flow through next. An architectural claim, encryption at rest, per-user isolation, who is structurally able to reach the content at all, is a narrower promise, and a narrower promise is one that's actually possible to verify and hold a company to.

This is the specific gap MemX (memx.app) is built to sit on the right side of. MemX stores what people actually want remembered: photos, documents, voice notes, WhatsApp messages, made searchable later. That's sensitive material by definition, closer to a health record or a bank statement than a throwaway chat message. MemX describes its privacy posture as private by architecture: encryption at rest, customer-managed encryption keys, and per-user isolation built into the storage layer, with on-device processing where the task allows it. What MemX is not building is a pipeline that routes personal memories to an outside contractor workforce, paid by the hour, to read and score them for model training. That's a structural choice about what the product is for, not a promise sitting in a policy document waiting to be tested by a lawsuit. It doesn't make MemX end-to-end encrypted or zero-knowledge, and MemX doesn't claim to be either. It means the privacy claim rests on what the system is built to do, not on a sentence that has to hold up entirely on its own.

The practical lesson isn't that AI chatbots are unsafe to use. It's that 'we don't share your data' and 'no outsider reads your data' are two different sentences, and it often takes a lawsuit to reveal which one a company was actually making. Reading the specific language, encryption at rest, per-user isolation, who is structurally able to reach your content and under what process, is a better five minutes than trusting the general tone of a privacy page.

Frequently Asked Questions
01Is the OpenAI Project Lily lawsuit real?

Yes. Vredenburgh v. OpenAI OpCo, LLC (case 3:26-cv-10527) was filed September 16, 2026 in the U.S. District Court for the Northern District of California and served on OpenAI September 21. It's a proposed class action, meaning allegations, not a finished case; OpenAI's response is due October 13, 2026.

02What does the Project Lily lawsuit accuse OpenAI of doing?

The complaint alleges OpenAI let outside contractors, hired through a staffing firm, read real ChatGPT conversations to summarize and score responses for model training, without disclosing that specific practice in its Privacy Policy or Terms of Service.

03Did OpenAI ever disclose that humans might read ChatGPT chats?

OpenAI's help center says authorized personnel and 'trusted service providers' may access content to improve model performance unless a user opts out. The lawsuit argues that general language differs from disclosing a named contractor workforce reading full conversations as routine work.

04How do I stop contractors from reading my ChatGPT conversations?

In ChatGPT, open Settings, then Data Controls, and turn off 'Improve the model for everyone', or select 'do not train on my content' in OpenAI's Privacy Portal. This affects new conversations going forward, not ones already reviewed.

05What does 'private by architecture' mean, and is it the same as end-to-end encryption?

No. It means encryption at rest, customer-managed keys, and per-user isolation are built into the system itself. End-to-end encryption and zero-knowledge are different, stronger claims meaning even the provider can't read content; 'private by architecture' doesn't claim that.

Was this article helpful?

Found this useful? Share it with someone who needs it.

Free · iOS, Android & WhatsApp

Stop losing what you save.
Let MemX remember it for you.

Every screenshot, photo, PDF and voice note: captured, encrypted, and instantly searchable. Ask in plain English, get the answer in seconds.

  • Reads text inside images and handwriting
  • Private and encrypted by default
  • Free to start, no credit card

Takes under a minute to set up. Your data stays yours.

Aditya Kumar Jha
Written by
Aditya Kumar JhaLinkedIn

Founding engineer at MemX, where he builds the website, backend, and data systems. Also a published author of six books on Amazon KDP, writing on AI, memory, and behavior.

Keep reading

More guides for AI-powered students.