The NSA, CISA, and the FBI issued a joint cybersecurity advisory on September 8, 2026, naming six AI companies, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, and describing them as systematically extracting outputs from major US AI models to train their own. This isn't a vague warning about competition from Chinese AI labs. It's a numbered federal advisory, AA26-251A, that names specific companies and specific model versions on the record.
Joint advisories from all three agencies together are uncommon, and naming individual companies by name, alongside individual model version numbers, is more specific than the general warnings about AI competition that have circulated in policy conversations over the past few years. That specificity is the reason this advisory is worth reading past the headline, not just noting that it exists.
The advisory describes activity dating back to at least late 2024, and puts a scale on it: billions of tokens extracted across millions of exchanges and requests against major US models. That's the government's characterization, published jointly by three agencies rather than floated by a single source.
What the advisory actually says
Six companies are named directly: DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI. The advisory ties the described activity back to at least late 2024, meaning the agencies are describing something they say has been going on for close to two years, not a single incident. The named scale, billions of tokens across millions of exchanges and requests, applies to the pattern across these companies against major US models.
- DeepSeek
- Moonshot AI
- Alibaba
- MiniMax
- StepFun
- Z.AI
All six are prominent AI developers that have released their own consumer or developer-facing models over the past two years. The advisory groups them together as sharing a pattern, extracting outputs from major US models at scale, even though it goes into the most specific detail on one of them, DeepSeek, by name.
What distillation means in this context
Distillation, in the ordinary machine learning sense, means training a new model on the input-output pairs produced by an existing one, using a rival's prompts and responses as training data instead of building from scratch. Done at the scale the advisory describes, billions of tokens across millions of requests, it lets a company skip a large share of the cost and time of training a frontier model from zero, while still ending up with something that behaves a lot like the model it learned from. The advisory frames this as the specific pattern under investigation, not ordinary API usage.
There's a legitimate version of this technique. Researchers distill smaller, cheaper models from larger ones openly and routinely, and plenty of that work is disclosed and unremarkable. What separates the legitimate version from what the advisory describes is scale, method, and disclosure: quietly running millions of queries against a rival's paid product specifically to harvest training data is a different act than publishing a paper about compressing your own model. The advisory's language, again, calls this the critical core of the named companies' strategy rather than a footnote, which is the part that turns a technical practice into a named security concern.
DeepSeek, specifically
DeepSeek is named as targeting GPT-4 and GPT-5 along with multiple Claude versions, using that extracted data to train its own R1 and V3 models. The advisory's own language is blunt about how central this is to DeepSeek's approach: it describes this kind of distillation as the critical core of the company's development strategy, not a supplement bolted onto an otherwise from-scratch training process.
That framing, critical core rather than supplement, is doing real work in the advisory's language. A supplement implies a company trains its own model from the ground up and uses extracted outputs to polish the result at the margins. A critical core implies the extracted data is load-bearing, that R1 and V3 as products would look meaningfully different, or cost meaningfully more to build, without it. The agencies chose the stronger of those two descriptions when they wrote the advisory, and named DeepSeek specifically rather than folding it into the general group of six.
Which models were allegedly targeted
The advisory lists specific model versions across four families. On the Claude side: 3.7, Sonnet 4 and 4.5, Opus 4.1, Haiku 4.5, and Fable 5. On the GPT side: 4, 4o, 4 Mini, 4 Nano, 5, and 5.5. On Gemini: 2, plus 2.5 Pro and Flash. And on Grok: 3 Mini and 4. Naming individual model versions, rather than speaking generally about US AI, is what makes this advisory unusually specific compared to earlier warnings about AI competition.
That breadth of named versions is itself informative. This isn't a claim that one older, possibly under-secured model got scraped once. Spanning nearly every current release across all four major US labs, from Claude's oldest listed version through its newest, from GPT-4 through GPT-5.5, from Gemini 2 through 2.5, and both current Grok variants, suggests the agencies are describing a sustained, evolving pattern that tracked each lab's releases as they shipped, rather than a single point-in-time incident against one product.
What the advisory doesn't tell us
It's worth being precise about what a joint advisory like this is and isn't. It's the assessment of three US government agencies, published under their own names with a formal tracking number, which carries real weight. It isn't an independently audited forensic report open to public review, and it doesn't walk through the underlying evidence for a general reader the way a court filing or an academic paper might. Treating the advisory as authoritative on what it states, while being clear-eyed that it represents one set of institutions' characterization rather than a neutral third-party verdict, is the fair way to read a document like this.
The advisory also doesn't specify, at least in what's summarized here, exactly how each of the five companies other than DeepSeek used the outputs it says they extracted, whether for training, evaluation, benchmarking against, or some other purpose. DeepSeek gets the specific R1 and V3 training claim; the other five are named as part of the same pattern without that level of individual detail. That's a real asymmetry in the public record, and it's why the comparison below treats DeepSeek differently from the rest rather than assuming identical activity across all six.
| Named Company | Models Allegedly Targeted | Primary Use (per advisory) |
|---|---|---|
| DeepSeek | GPT-4 and GPT-5, multiple Claude versions | Training its own R1 and V3 models; advisory calls this distillation the critical core of its development strategy |
| Moonshot AI | Named in advisory; distillation activity since late 2024 | Named in advisory; distillation activity since late 2024 |
| Alibaba | Named in advisory; distillation activity since late 2024 | Named in advisory; distillation activity since late 2024 |
| MiniMax | Named in advisory; distillation activity since late 2024 | Named in advisory; distillation activity since late 2024 |
| StepFun | Named in advisory; distillation activity since late 2024 | Named in advisory; distillation activity since late 2024 |
| Z.AI | Named in advisory; distillation activity since late 2024 | Named in advisory; distillation activity since late 2024 |
Why the scale figure matters
Billions of tokens across millions of exchanges and requests is not a number that happens by accident or through a handful of curious researchers testing a competitor's product. Reaching that volume requires sustained, automated querying over an extended period, exactly the kind of activity the advisory says has been running since at least late 2024. The scale is the detail that turns this from an interesting technical footnote into the subject of a joint federal advisory: it describes something closer to an industrial-scale data collection operation than isolated instances of a developer poking at a rival's API.
What happens next is still open
A joint advisory is a starting point, not a conclusion. What follows, whether that's a policy response, changes at the named companies, or nothing visible at all, hasn't been determined as of September 2026. What's already certain is narrower and more immediate: the advisory exists, it names six companies and dozens of specific model versions, and it's now part of the public record that anyone evaluating these AI products has to weigh.
Why this matters if you're just a user, not a security team
None of this advisory is addressed to individual chatbot users, it reads like a document written for network defenders and policy staff. But it lands on a fact that's easy to miss day to day: the model you're chatting with tonight isn't a fixed thing. It's a moving target, both in the sense that labs ship new versions on their own schedule, and now, per this advisory, in the sense that its own outputs may themselves be a target other companies are mining at scale to build competing products.
That's a reason to separate two questions that usually get treated as one: which AI model is best this week, and where does your own history with these tools actually live. Those don't have to be the same answer, and treating them as identical means every time the first answer changes, you risk losing the second.
Consider what actually happens when a person switches from one AI assistant to another, which is already common and, per the advisory's framing of a fast-moving competitive field, likely to keep happening. Months or years of context, how you like things explained, what projects you're working on, what you've already told the model so you don't have to repeat it, typically stay locked inside whichever chat interface holds them. None of that is about distillation or national security. It's a much more ordinary problem that this advisory happens to put back in the spotlight: your own history with an AI tool is only as durable as your commitment to one specific product.
Your memory shouldn't depend on which model wins
If the model you're using today might get retrained, distilled, or overtaken by something built partly on someone else's extracted data tomorrow, tying your own memory to that one vendor means your history moves only as far as that vendor's fortunes do. MemX's position is that your memory of your own conversations should be portable across whichever model you end up using, ChatGPT, Claude, Gemini, or whatever replaces them, and private by architecture rather than locked inside a single lab's product. The churn in which company looks trustworthy this month is real. It shouldn't be the thing that decides whether you keep your own context.
None of that requires taking a side in the dispute this advisory describes. Whether the named companies are found to have acted as alleged is a separate question from the more basic one this whole episode surfaces for any regular user: the AI landscape is shifting fast enough, between new model releases, competitive pressure, and now advisories like this one, that betting your entire memory and context on a single vendor staying exactly where it is today looks increasingly like the wrong default.
If you use more than one AI assistant, or expect to switch as models change, check whether your memory and context are exportable before you need to move them, not after.
01What is CISA advisory AA26-251A about?
A joint September 8, 2026 advisory from the NSA, CISA, and FBI naming six companies, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, for allegedly extracting outputs from major US AI models at scale since late 2024.
02Did the US government say DeepSeek stole from Claude and GPT?
The advisory says DeepSeek targeted GPT-4, GPT-5, and multiple Claude versions to train its own R1 and V3 models, and calls this distillation the critical core of DeepSeek's development strategy.
03What does AI model distillation mean?
Training a new model on another model's prompts and outputs at scale, instead of training from scratch, so the new model behaves similarly at a fraction of the original training cost.
04Which AI models does the advisory say were targeted?
Claude 3.7, Sonnet 4 and 4.5, Opus 4.1, Haiku 4.5, and Fable 5; GPT 4, 4o, 4 Mini, 4 Nano, 5, and 5.5; Gemini 2 and 2.5 Pro and Flash; and Grok 3 Mini and 4.
05Is my data at risk if I use ChatGPT, Claude, or Gemini?
The advisory concerns companies allegedly extracting model outputs at scale, not individual user accounts. It's still a reason to think about keeping your own conversation history independent of any single AI vendor.
A joint advisory naming six companies and specific model versions, backed by the NSA, CISA, and the FBI together, is a rare, specific document, not the usual background noise about AI competition. It names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI on the record, ties the described activity back to at least late 2024, and puts a number on it, billions of tokens across millions of exchanges, that's hard to read as anything other than sustained and deliberate.
Whatever happens next between these labs and the US government, whether through policy responses, restricted access, or nothing visible at all, the practical takeaway for anyone using these tools daily is smaller and more useful: the model underneath you can change for reasons that have nothing to do with you, and your own memory of your own conversations is worth keeping somewhere that doesn't rise or fall with any single one of them.
