Yes, ChatGPT chats can be kept, and a July 2026 court fight made that concrete: roughly 20 million ChatGPT conversations were preserved and produced to a federal court as evidence. The openai 20 million chatgpt conversations sample sits at the center of a copyright case brought by the New York Times and other news outlets, and it is a plain reminder that pressing delete on a chat does not guarantee the chat is gone. If a court tells a company to hold onto its data, your conversations can outlive the moment you thought you erased them.
The short version: those conversations survived because a court ordered them preserved, not because anyone forgot to hit delete. Here is what happened in the case, why so many chats existed in the first place, and what the answer to "are chatgpt chats kept" means for a normal person who just wants a private place to store their own information.
What happened: the July 2026 sanctions fight
On July 9, 2026, the New York Times and a group of co-plaintiffs asked a federal judge to sanction OpenAI over how it handled evidence in their copyright lawsuit. The other outlets in the group include the New York Daily News, the Center for Investigative Reporting, the Chicago Tribune, and Ziff Davis, among other newspaper groups. Their motion accuses OpenAI of discovery misconduct: for more than two years, they say, OpenAI told the court it could not search its training data and ChatGPT output logs for their articles, even though it had built tools to run exactly those searches.
One specific remedy the publishers asked for: bar OpenAI from relying at trial on a 20-million-conversation sample of ChatGPT logs. According to reporting on the filing, OpenAI produced that sample to the court in December, down from an original request for 120 million, and the court found the sample unusable because of heavy redactions. The publishers say the Times alone has spent more than 28 million dollars fighting AI companies in court so far.
The publishers built their argument on testimony from an April deposition. According to TechCrunch's account, an OpenAI data privacy engineer described internal searches of the company's training corpus for copyrighted journalism, and a database of roughly 78 million de-identified ChatGPT conversations that OpenAI kept to gauge how often its models reproduced protected content. The publishers also pointed to a tool set they call Project Giraffe, which they say used a filtering method to detect and record when ChatGPT outputs echoed copyrighted text, deployed soon after the lawsuit was filed. OpenAI frames these as ordinary, good-faith engineering steps rather than concealment.
The headline number to remember: about 20 million ChatGPT conversations were retained and handed to a court as evidence. These were not chats users chose to publish. They were ordinary conversations sitting in OpenAI's systems, then pulled into a legal case.
OpenAI denies wrongdoing. A company spokesperson called the allegations false and argued that the publishers are trying to reach into private user data as their own case weakens. Both sides agree on one underlying fact, though: the conversations existed, in volume, and were available to be searched and produced.
Why did 20 million chats still exist?
They existed because a court preservation order overrode the usual cycle of deletion. When a company is sued, it typically comes under a legal hold: a duty to retain any data that could be relevant to the case. That hold sits on top of, and can override, a provider's normal retention and deletion settings. So a chat you would expect to age out or disappear can instead be frozen in place for the length of the dispute.
In this case the publishers went further and alleged that OpenAI deleted or compressed large volumes of ChatGPT conversations, making them harder to search, despite the order to preserve that data. OpenAI disputes that characterization. The point for a reader is not who wins the argument. It is that the retention of your chats was being decided by a court and by lawyers, not by the delete button in your account.
Legal holds work this way in almost every industry, not just AI. Once a company reasonably anticipates litigation, it is expected to suspend routine deletion for anything relevant and preserve it. Email, chat logs, and product data all get frozen. The difference with a consumer AI service is scale and intimacy: the frozen material is not corporate memos, it is millions of ordinary conversations that individual people typed into a box, often assuming the exchange was fleeting. When the hold lands, those exchanges become records to be counted, sampled, and argued over.
When any online service is under litigation or a regulatory hold, its published deletion policy is not the final word. A legal hold can require the company to keep data it would otherwise remove, and it usually will not notify individual users that their content is affected.
Does deleting a chat actually delete it?
Not always, and not immediately. Deleting a conversation almost always removes it from what you can see. What happens on the company's side depends on three things: the provider's retention policy, its backup and logging systems, and any legal obligation that forces it to hold the data. The visible chat can vanish while a copy lives on in logs or backups for a period the provider defines, or that a court extends.
Most consumer AI chatbots also run on a shared model: your conversations sit inside a large multi-tenant system, and by default many services may retain them to improve the product or to meet safety and legal requirements. That is the setup that made a 20-million-conversation sample possible to produce in the first place. The data was centralized, searchable, and under the company's control rather than yours.
The word that matters: discoverable
In a lawsuit, evidence that a party holds and that is relevant to the case is discoverable, meaning the other side can demand it and a judge can order it produced. If a company keeps your conversations, and those conversations become relevant to some dispute, they can be swept into discovery. You are usually not a party to that case, and you often will not know it happened. Anonymization helps, but the July 2026 fight shows how contested even de-identified samples can be once they reach a courtroom.
The privacy tension in this case runs both ways, which is what makes it worth watching. OpenAI has publicly argued that handing over millions of conversations would expose private user data, and it resisted the broader demands on those grounds. The publishers counter that they need the logs to prove infringement and that OpenAI's own conduct created the problem. A normal user sits in the middle of that fight without a vote, which is the clearest reason to be deliberate about which conversations you feed into a shared system at all.
What this means for your own information
The lesson is not that AI tools are unsafe to use. It is that where your data lives, and who controls it, changes what can happen to it later. A conversation stored in a shared cloud system, retained by default and under one company's control, carries a different risk profile than information kept in storage that is isolated to you and shaped by settings you choose.
For a chat you send to a general assistant, retention is the provider's call. For your own documents, photos, voice notes, and messages, you can pick tools that put you in charge of what is kept and for how long. The distinction below is about control and design, not about any single company being good or bad.
Two design choices drive most of the risk. The first is pooling: when every user's data flows into one shared system, that system becomes a single, searchable target for discovery, subpoenas, or a breach. The second is default retention: when a service keeps everything unless you dig into settings to say otherwise, the volume of stored data grows quietly, and so does the amount that can be produced later. The 20-million-conversation sample was possible precisely because both choices pointed toward keeping more, in one place, under the provider's control.
| Question | Typical AI chatbot | Private-by-architecture app |
|---|---|---|
| Where do your chats live? | In a large shared cloud system run by the provider | In storage isolated to your account |
| Who decides retention? | The provider's policy, plus any legal hold | You decide what is kept and what is removed |
| Used to train models? | Often yes by default, unless you opt out | No training on your data |
| Encryption of stored data? | Varies by provider | Encryption at rest, with customer-managed keys |
| Can a legal hold freeze your data? | Yes, and you may never be told | Data isolation limits exposure, but no service is beyond the law |
The honest version: no consumer product can promise your data is beyond the reach of a valid court order. What a well-designed product can do is keep your information isolated to you, avoid training on it, and give you the controls to decide what is stored. That is a real difference, even if it is not a magic shield.
How MemX approaches the same problem
MemX is a consumer memory app: it stores your documents, photos, voice notes, and messages so you can ask questions and get instant answers with the source. It is built to be private by architecture. That means per-user data isolation so your content is not pooled into a shared training set, encryption at rest with customer-managed keys, on-device processing where possible, and no training on your data. You decide what goes in and what comes out. MemX is not end-to-end encrypted and is not zero-knowledge, and no app can place your data outside the law, but the design keeps your information yours by default rather than centralized and searchable by the provider.
What you can do right now
- Turn off chat history or training in any AI tool you use, and check whether deletion is immediate or delayed by a retention window.
- Assume anything you type into a shared cloud assistant could, in theory, be retained and later requested in a legal process.
- Keep sensitive personal records in storage that is isolated to your account, not in a general chatbot's history.
- Read the retention section of a service's privacy policy, not just the marketing page, so you know how long deleted items actually persist.
- Favor tools that state plainly that they do not train on your data and that let you control what is kept.
Frequently asked questions
01Are ChatGPT chats kept after you delete them?
They are removed from your view, but retention on the company's side depends on its policy, its backups, and any legal hold. A court preservation order can require a provider to keep conversations it would otherwise delete, so deletion is not always permanent.
02What is the OpenAI 20 million ChatGPT conversations case about?
In July 2026, the New York Times and other news outlets sought sanctions against OpenAI in a copyright lawsuit. Part of the fight involved a roughly 20-million-conversation sample of ChatGPT logs that OpenAI produced to the court and that the publishers asked the judge to exclude.
03Can my private chats become court evidence?
If a company retains your conversations and they are relevant to a lawsuit, they can be discoverable, meaning a court can order them produced. You are often not a party to the case and may never be notified that your data was involved.
04Does a court order override an app's delete button?
Yes. A legal hold sits on top of a provider's normal retention settings and can require the company to preserve data it would otherwise remove. That preservation duty can last for the length of the dispute.
05How can I keep personal information more private?
Use tools that isolate your data to your account, do not train on it, and let you control retention. Store sensitive records outside general chatbots, and check each service's actual retention window rather than trusting the delete label alone.
