A Chinese-speaking hacking group built an AI workflow with persistent memory across an entire vulnerability-hunting campaign, and it worked well enough to surface more than a dozen possible zero-day findings in a single month. Anthropic disclosed the operation, tracked as GTG-10007, in a threat intelligence report published September 10, 2026, and the company says it detected and disrupted the cluster rather than let it keep running.
What Anthropic's September 2026 Report Covers
Anthropic's report, titled Detecting and countering misuse of AI: September 2026, covers misuse cases the company disrupted between December 2025 and August 2026. The cases span Claude's Haiku, Sonnet, and Opus model families. Anthropic states it found no misuse on its Fable or Mythos model families over the same period, a detail worth including because it shows the report is specific about where the activity was found rather than treating every model the same way.
Why Anthropic Publishes These Reports
Reports like this one exist because frontier AI labs sit in an unusual position. They can see when their own models are being used for harm in ways that outside security vendors generally cannot, since the misuse runs through infrastructure the lab itself operates. Publishing a detailed account, naming clusters, sectors, and techniques, serves two purposes at once. It gives defenders concrete indicators to watch for, and it gives the public a way to judge whether a lab is actually finding and stopping misuse or merely asserting that it takes safety seriously. Anthropic's September 2026 report falls into the first category: a fairly granular account of three unrelated clusters, rather than a vague assurance.
That kind of granularity is also what makes a report like this useful to read skeptically rather than take on faith. Naming a specific cluster, a specific likely location, a specific count of targeted organizations, and a specific outcome, more than a dozen possible zero-day findings in a single month, gives outside researchers something concrete to weigh, rather than a marketing claim that resists being checked at all. Whether every detail eventually holds up under further scrutiny is a separate question from whether the report was written in a way that invites scrutiny in the first place.
Inside GTG-10007: A Persistent-Memory Vulnerability Hunter
The operators behind GTG-10007 are assessed by Anthropic as Chinese-speaking and likely based in Changsha, in Hunan province. They built an autonomous vulnerability-research workflow on Claude and pointed it at roughly 50 organizations spanning a wide range of sectors.
- Education
- Retail
- Energy
- Technology
- Healthcare
- Finance
- Manufacturing
- Multiple government agencies globally
All eight sectors named above are prominent in the report's account of GTG-10007's targeting, though the actual public sector-by-sector breakdown of which organizations belonged to which category has not been disclosed. What Anthropic has disclosed is the aggregate scope, roughly 50 organizations, and the sectors those organizations span, which is enough to establish that this was a broad, opportunistic campaign rather than a narrowly targeted one against a single industry.
How an Autonomous Vulnerability-Research Workflow Operates
Vulnerability research, whether done by a person or by an AI-driven workflow, generally follows the same rough loop: examine a piece of software or firmware, form a hypothesis about where it might fail, test that hypothesis, record what happened, and use the result to decide what to try next. Done manually, that loop can take a skilled researcher days or weeks per target, because each step depends on carefully understanding what the previous step revealed. An autonomous workflow built on a language model can run that loop far faster, generating hypotheses, writing test code, and interpreting results with far less human involvement at each step. What GTG-10007 added on top of that loop was continuity between sessions, so the workflow did not have to re-derive its understanding of a given piece of firmware every time it picked the work back up.
Network appliance firmware, the specific target named in Anthropic's report, is a particularly attractive fit for this kind of sustained analysis. It tends to receive less scrutiny than mainstream operating systems or applications, changes less often once deployed, and often sits at the edge of a network with broad visibility into the traffic passing through it. A workflow patient enough to study the same firmware repeatedly, across many sessions, is well suited to exactly that kind of target, which is part of why the combination of an already fast automated loop plus memory that carries forward was enough to surface more than a dozen possible zero-day findings in a single month.
Why Persistent Memory Made This Workflow More Effective
A one-off prompt against a piece of firmware gets you a single pass of analysis, and then the context resets. Ask the same question tomorrow and the model starts over, with no memory of which paths it already ruled out. A workflow with persistent memory skips that reset. Anthropic's report describes the operation as maintaining persistent campaign memory: target lists, harvested credentials, engagement state, and standing instructions were saved and carried forward across sessions, so the workflow could pick up exactly where it left off.
Think of the difference between a support agent who reads your entire ticket history before responding and one who greets you fresh every time and asks you to re-explain the problem from the start. The first is faster and more useful precisely because nothing has to be repeated. Persistent memory gives an AI workflow the same advantage over a series of disconnected prompts. Applied to customer support, that advantage is a convenience. Applied to a month-long campaign against firmware inside fifty organizations, the same advantage is what turns a string of individual attempts into a working, cumulative campaign.
That is not a special property of malicious use. It is the same reason persistent memory is useful for a security researcher's own legitimate work, or for any long-running technical project that spans weeks instead of a single sitting. GTG-10007 is best understood as that ordinary property of memory being pointed at other people's infrastructure instead of the operator's own project.
Detected and Disrupted, Not Ongoing
Anthropic frames GTG-10007 as a misuse case it found and shut down, not a live, unaddressed threat sitting out in the open today. The report is explicitly about detecting and countering misuse, and its account of this cluster covers activity through August 2026. That distinction is worth sitting with for a moment: this is a record of a threat actor's workflow being caught, not a warning that it is still running unchecked as of September 2026.
What Organizations Defending Infrastructure Should Take From This
The defensive value of this report is not the exact identity of GTG-10007, since Anthropic's disclosure does not name the affected organizations. It is the shape of the technique: a persistent, memory-driven workflow iterating on network appliance firmware over an extended period rather than in a single scan. Keeping firmware on network appliances patched and monitored, treating vendor security advisories for edge devices as high priority rather than routine, and assuming a sufficiently motivated actor can now sustain analysis over weeks rather than hours are all more relevant in light of a report like this one. None of those practices are new. What is new is a documented case of an AI-driven workflow with the patience and continuity to make slow, iterative firmware analysis practical for an attacker at this scale.
Two More Clusters From the Same Report
GTG-10007 was not the only cluster Anthropic disclosed in this report. A Russian-linked group, tracked as GTG-20006, targeted more than 20 organizations concentrated in Ukraine and Europe, including government ministries, defense and intelligence bodies, and diplomatic missions, with activity extending to the Middle East and parts of Asia. Separately, a group affiliated with the ShinyHunters collective, tracked as GTG-50014, was tied to large-scale financially motivated data theft. Three distinct clusters, three different sets of operators, disclosed in the same reporting window.
It is worth pausing on how different these three clusters are from each other in purpose. GTG-10007 was built around technical vulnerability discovery. GTG-20006 was aimed at government, defense, and diplomatic entities tied to a specific geopolitical conflict. GTG-50014 was about large-scale data theft rather than vulnerability research at all. Grouping them only by the fact that Anthropic disrupted all three in the same window shows the breadth of what one company's misuse-detection effort has to cover, not a single unified threat with one motive behind it.
It is worth being precise about what carries over between the three and what does not. Persistent, cross-session memory is the detail specific to GTG-10007, tied directly to the firmware-hunting workflow described in the report. Nothing in Anthropic's disclosure says GTG-20006 or GTG-50014 used the same technique. The pattern across all three is not a shared method, it is that three separate actors, with three separate goals, were each running Claude-based workflows that Anthropic's own systems caught within the same reporting window.
| Disclosed Cluster | Attributed To | What It Targeted |
|---|---|---|
| GTG-10007 | Chinese-speaking operators, assessed as likely based in Changsha, Hunan province | Roughly 50 organizations across education, retail, energy, technology, healthcare, finance, manufacturing, and government |
| GTG-20006 | Russian-linked group | More than 20 organizations, concentrated in Ukraine and Europe: government, defense, and diplomatic targets, extending to the Middle East and Asia |
| GTG-50014 | Group affiliated with ShinyHunters | Large-scale, financially motivated data theft |
Not Just Chatbots: Why This Report Matters Beyond Claude
GTG-10007's workflow is a specific example of a broader shift already underway across the industry: AI systems moving from answering questions to taking sustained, multi-step actions on someone's behalf. Persistent memory is the feature that makes that shift work at all, whether the workflow is hunting for firmware vulnerabilities, managing a long research project, or handling a customer support queue. Anthropic catching this particular misuse case does not mean the underlying technique, memory-driven autonomous workflows, is confined to one company's models. It means one company's detection systems caught one case built on its own infrastructure, which is a narrower and more specific claim.
The Double-Edged Nature of AI Memory
Here is the tension worth naming directly instead of glossing over. The exact property that makes persistent memory valuable for honest, everyday use, not having to re-explain your context every session, is the same property that made GTG-10007's workflow effective at hunting vulnerabilities across a month-long campaign. Memory does not know whether the work it is carrying forward is legitimate. It just carries context forward, faithfully, regardless of who is directing it.
That is a real, uncomfortable fact about the technology, and it is worth sitting with rather than explaining away. It is not, on its own, a reason to conclude that memory is unsafe by nature. The same report that describes what GTG-10007 built also describes Anthropic catching it, which is the part of the story that a purely alarmed reading of these facts tends to skip.
This tension is not unique to Claude or to Anthropic. Any AI system that offers persistent memory as a feature, for coding, for research, for personal use, is offering the same double-edged capability described in this report, whether or not it has been used this way yet. The reasonable response to a report like this one is not to conclude that persistent memory should not exist. It is to take seriously that the property is real, that it cuts both ways, and that detection and disclosure, the parts of this story that led to GTG-10007 being caught, matter as much as the underlying capability itself.
The meaningful difference is not whether memory persists. It is who controls it and what it is pointed at. An autonomous vulnerability-research workflow's memory is built to serve an operator's campaign against a target that never consented to being studied. MemX's persistent memory is built around the opposite premise: one person's own continuity across their own ChatGPT, Claude, and Gemini conversations, controlled by that person, and private by architecture. That is a distinction about design intent and control, not a claim that MemX or any product prevents the kind of misuse Anthropic describes in this report.
When evaluating any AI tool with persistent or cross-session memory, the useful question is not just what it remembers. It is who can direct that memory and toward what end, whether you are a security team reviewing an agentic workflow or an individual choosing a personal AI assistant.
Detection, disruption, and public disclosure are the pattern worth tracking as AI-driven workflows keep getting more capable. The technique described in this report will not disappear because one cluster got caught. The open question going forward is whether detection keeps pace with capability, and reports like this one are among the only public evidence pointing either way.
01What is GTG-10007?
GTG-10007 is a Chinese-speaking hacking cluster Anthropic disclosed in its September 2026 threat intelligence report. It built an autonomous vulnerability-research workflow on Claude with persistent memory, targeting roughly 50 organizations.
02Did hackers use Claude to find zero-day vulnerabilities?
Anthropic's September 2026 report says a workflow built on Claude, tracked as GTG-10007, surfaced more than a dozen possible zero-day findings in a single month before Anthropic detected and disrupted the cluster.
03Is the GTG-10007 hacking campaign still active?
No. Anthropic describes GTG-10007 as a case it detected and countered, covering activity through August 2026, not an ongoing, unaddressed threat as of the report's publication.
04What other AI misuse did Anthropic report in September 2026?
The same report disclosed GTG-20006, a Russian-linked group targeting government, defense, and diplomatic entities concentrated in Ukraine and Europe, and GTG-50014, a ShinyHunters-affiliated group tied to large-scale data theft.
05Why does persistent memory make an AI misuse workflow more effective?
It lets a workflow build on prior sessions instead of restarting analysis from zero each time, the same property that makes persistent memory useful for legitimate long-running work.
