On September 10, Anthropic published its September 2026 threat intelligence report, summarizing misuse of Claude in areas such as cyberattacks and state surveillance. The report covers activity across seven areas that the company disrupted between December 2025 and August 2026. It shows AI taking on sustained work for attackers, and it describes a case in which a surveillance system developed with Claude kept running on local models after access was cut off. The question it raises: how far do measures that detect misuse and suspend accounts reach once the output has left the AI's hands?
The shift from advice to execution had already begun
This is not the first report to describe Claude acting as the one carrying out an attack. In its August 2025 report, Anthropic disclosed intrusions and data extortion carried out using Claude Code. In November of the same year, it announced that an attacker it assessed to be a Chinese state-sponsored group had attempted to breach about 30 organizations and succeeded in a small number of cases.
In the November case, AI reportedly handled 80 to 90 percent of the work in the operation. That figure, however, is Anthropic's own assessment of that operation. The AI also fabricated credentials and mistook public information for confidential data. Doing a lot of the work is not the same as getting it right.
What has changed this time is that this style of operation has spread across the range of attackers under study. Besides state-linked groups, financially motivated criminals and politically motivated individuals were also assigning tasks to multiple AI agents. Humans choose targets, AI carries out reconnaissance and intrusion, and humans review the results. Anthropic's observation is that use has moved from getting advice in a chat to delegating the execution and management of work.
This is not, however, a statistic covering all misuse. Anthropic says it selected new cases that merit particular attention rather than typical usage. The models that appear are mainly in the Haiku, Sonnet, and Opus families; Fable- and Mythos-class models are not included, apart from one case involving distillation. Because usage volumes and access conditions are not comparable across models, these differences cannot be tied directly to how well safeguards work.
Other companies' observations differ in timing and scope as well. In its February 2026 threat report, Google said AI was boosting attacker productivity in late 2025, but it had not confirmed that state-linked or influence-operation actors had gained capabilities that fundamentally change the threat landscape. Anthropic's new report cannot be read as an industry-wide autonomy rate or growth rate in damage.
Less effort per attack, and stolen AI access powering the next one
In "GTG-20006," which Anthropic assesses to be Russia-linked espionage, Claude was used for everything from reconnaissance and intrusion support to processing stolen information. A mechanism was also built so that when security products detected the attack tools, the AI repeatedly modified and rebuilt them. Humans stayed involved in adjusting operating policy and work procedures while handing the hands-on work to the AI.
When defenders add a new detection, attackers have to rewrite their tools. If AI handles that cycle, the time gained from each detection may shrink. Alongside the performance race over whether AI can discover unknown vulnerabilities, the ability to chain together existing intrusion methods and absorb differences between targets is becoming important.
In financially motivated attacks too, Anthropic describes a case in which a stolen developer token gave attackers cloud administrative access at a victim organization in about three hours. That is not how long attacks generally take, but it concretely shows the danger of a single privilege serving as a foothold to spread damage quickly.
There is also a link to Microsoft's July 31 "CaptiveCrunch" investigation, which concerns intrusions via hotel and similar networks. From early May, Microsoft observed activity in which traffic on networks with captive authentication pages was manipulated to steer users to fake update notices and credential prompts. It assessed the actor to be Storm-2945, a Midnight Blizzard activity cluster, and said AI supported a substantial part of the operation.
Anthropic likewise says its attribution of GTG-20006 is consistent with public reports tying the activity to Midnight Blizzard. Microsoft, however, thanks Anthropic and OpenAI for assisting its investigation. So while this is a technical observation by a separate company, it is not a fully independent replication in terms of information sources, nor a separate measurement of how much AI contributed.
Access to the AI that drives attacks was itself being stolen and resold. API keys that customers had exposed in apps or public code, along with session tokens showing logged-in status, were collected, and intermediaries passed them on to other users. Attackers then used that access as the computing resource for their next attack.
In the ShinyHunters-related case, the stolen keys leaked from customer environments; Anthropic's own systems were not breached. Even without an intrusion into the company, a leaked customer key hands the right to use the model to the criminal side. AI API keys are now something to protect for both billing management and attack prevention.
Seven areas, each with a different answer to "what did the AI do?"
Regarding the five cases in the biological research section, Anthropic does not claim the researchers had harmful intent. That is because research that helps treatment and research that can be diverted to dangerous uses overlap. "We found assistance that could be used for biological weapons" and "a biological weapon was completed" are very far apart.
The same distinction is needed in other areas. Selecting representative cases from each chapter of the report and matching what Claude did with the confirmed results and limitations gives the following.
| Area | What Claude did | Limits on reading the result |
|---|---|---|
| Cyberattacks | Supported reconnaissance and intrusion, processed stolen information, automated tool modification after detection | Humans were involved in target selection and checking results. Suspending an account does not mean victims have recovered |
| State surveillance | Designed and developed a surveillance system for Mali | The system in operation ran on local models. Cutting off Claude did not affect the deployed system |
| Influence operations | Created posts and fake personas, and developed the mechanisms for running accounts | Most of the content found drew little response from real users |
| Conventional weapons | Developed guidance software and a simulation environment | A live-fire rocket test in Yemen appears to have failed, and there is no evidence of successful operational deployment |
| Biological research | Drafted research plans and applications, and supported analysis | Not a report confirming harmful intent by researchers or the completion of a weapon |
| Fraud | Developed dating apps and conducted conversations posing as real people | The number of people engaged is not the number confirmed to have suffered financial loss |
| Unauthorized distillation | The company alleges outputs were diverted to train other models and for similar purposes | A matter of model-capability diversion and terms-of-service violations; it cannot be measured on the same scale as physical harm from weapons |
Comparing the seven areas by what the AI did and what results can be confirmed, a suspended account and a halted operation need to be distinguished, and Mali's already-running surveillance system is the explicit example.
What is compared here is the "task performed" and the "stage of the result" recorded in the same report covering activity from December 2025 to August 2026. Because the time periods and the base numbers of victims differ by case, this is not a ranking of danger. Separating the stage where AI builds software, the stage where actual activity is run, and the stage where physical results emerge also changes where an access cutoff can be effective.
In biological research, the avian influenza case was at an early planning stage, and the models available were Sonnet 4 and Haiku 4.5. Anthropic assesses the help they provided as limited, such as in research conception and data analysis. In a separate research application, however, Opus 5 helped draft the application in about an hour. That is the time taken to prepare the paperwork, not to finish an experiment.
In influence operations too, there was an example of a mechanism aimed at Malaysia that managed over 1,000 fake X accounts and called for 1 million artificial views, but that does not mean the target was reached. Most of the posts the company found drew almost no response from real users. Content distributed through existing state media reached a wider audience, so being able to mass-produce posts cheaply does not guarantee winning readers.
In the fraud case, across more than 20 dating apps, over 4,700 AI-generated personas conversed with at least 25,000 people during two weeks in April 2026. The operators also mixed in real humans, making it easier for targets to believe the other party was genuine. Looking only at AI's conversational ability makes it hard to grasp the deception that works across the app as a whole.
Distillation, meanwhile, refers to the technique of using one model's outputs to train another, and it is not always illegitimate. The report is concerned with unauthorized use and circumvention of access restrictions. It also describes a case in which an intermediary stored conversations with users and sold them to other developers. For companies, that raises not only a dispute over model intellectual property but also the question of where their own inputs end up.
Banning the account doesn't remove the surveillance system
"Lakana 360," which Anthropic reports was built for Mali's intelligence agency, was a surveillance system covering three domestic telecom operators and about 25 million SIM cards. The company explains that a single Claude user, apparently an independent local contractor working with the intelligence agency, delegated the main development work to the AI. The 25 million figure counts SIMs; it should not be read as 25 million citizens individually monitored.
In this case, Claude was not the production model analyzing surveillance records. It handled the design and development of the software, and the deployed system ran on in-house equipment and local models. Anthropic states that the account suspension prevented it from supporting development and design, but did not affect the already deployed system.
Revoking access does not erase software that has already been built, because the AI service used during development and the computing environment used in operation are separate. This shows that when reading the word "prevented" for misuse, one must also check where the output ended up. If the environment that continues the surveillance is outside the cloud, a response different from cutting off the cloud side is required.
In the biological research intermediary service, there were also moves to continue activity by switching connection destinations. Anthropic notes that after it suspended accounts connected to chikungunya virus research, the intermediary secured access again within days. In addition, a mechanism had been built to send questions Claude refused to other companies' models with looser restrictions.
Squeezing the moment the company detects and stops activity, and the activity that continues afterward, into a single success or failure obscures this difference. In Mali the output remained outside; at the intermediary service access was restored via another route. The former requires a response to an already running system, the latter requires tracking across connection routes.
Beyond models that refuse: managing the routes of use
On biological research, Anthropic recommends that, in addition to filters that identify dangerous content, mechanisms to verify individuals and their affiliated organizations, and records for investigating misuse, are needed. The wording of a question alone cannot reliably distinguish research aimed at treatment from research aimed at harmful uses. The company emphasizes providing advanced capabilities to users whose trustworthiness has been verified.
This demands difficult operational judgments from AI providers. What should be recorded to link suspicious use while protecting the confidentiality of research and business activity? When use passes through multiple accounts or intermediary services, how far can it be verified whose use it is? The report shows the distance between the test of whether a model can refuse a single question and the practical task of curbing misuse across the entire route of use.
Companies that use these services also have concrete reasons to review their existing measures. A leaked API key can be an entry point to the victim company's data and also usage rights to AI that attackers can use. If an intermediary service is used, one needs to confirm which model is actually provided at the connection destination and who stores the inputs and outputs. Even if one believes one has obtained cheap access, without visibility into the distribution route, the whereabouts of one's own information and usage rights cannot be traced.
Authentication and endpoint defenses do not become unnecessary because of AI. In response to CaptiveCrunch, Microsoft recommends combining passkeys and multifactor authentication with conditional access, and blocking unneeded device-code authentication. Not running software updates that a connection page at a travel destination suddenly demands is another measure. Precisely because attackers' work has become faster, reducing the entry points where credentials are handed over matters.
To evaluate future measures, it will be necessary to track not only the number of accounts suspended but also whether use has resumed via other routes and whether generated systems continue running externally. If AI providers, authentication platforms, and frontline defenders can connect that information, measures that cut off support for development could be linked to responses against the operations that cause harm.
