Rep. Ro Khanna has sent letters to OpenAI, Anthropic, Google, Meta, and SpaceXAI asking how they protect AI model "weights" from unauthorized access and theft, including from China. Reuters reported the news on October 1, 2026. On April 29, Senators Jim Banks and Chuck Grassley had also asked the CEOs of nine AI companies to respond in writing about how they protect model weights, among other topics. Weights are the collection of numerical values (parameters) an AI adjusts during training, and they largely determine a model's capabilities. What gets stolen is fundamentally different from a typical data breach, in which customer information or internal documents leak. So what happens when the weights of a model a company has fine-tuned itself are leaked? And are security measures designed for frontier AI developers also effective for ordinary companies?

AD

When weights leak, the trained model itself is stolen

In its 2024 report "Securing AI Model Weights," the U.S. think tank RAND Corporation defines weights as "learnable parameters that encode the core intelligence of an AI." The results of training that consumed vast computing resources and time are stored in files as numerical data. Load them into a compatible inference environment, and they run as a trained model.

In a typical data breach, what is taken is information such as customer lists or internal documents. When weights leak, by contrast, the trained model itself falls into a third party's hands.

In a research note published on May 17, 2026, the Cloud Security Alliance (CSA), a cloud security industry group, points out that weights can be copied in a matter of hours with a high-speed connection, and that an attacker with sufficient GPU infrastructure could start serving a stolen model within days without any additional training investment. However, this timeframe has not been confirmed by independent verification; it is CSA's estimate.

The same research note also warns that once a foundation model's weights leak, results equivalent to years of research and development and hundreds of millions of dollars in compute costs could pass to attackers who bore none of the development costs.

The letters from U.S. senators also take issue with the value of weights. Senators Banks and Grassley described the transfer of AI model weights to China as closer to the theft of a "finished AI product" than the theft of a blueprint.

AI model "distillation" works differently. It involves sending large numbers of queries to another company's AI model through its API and using the outputs to train one's own model. It can mimic a model's behavior without breaking into the other party's servers, but it does not yield the original model's weights themselves.

In fact, on February 23, 2026, Anthropic announced that three Chinese companies, DeepSeek, Moonshot, and MiniMax, had used about 24,000 fraudulent accounts to carry out more than 16 million exchanges with Claude in an attempt at distillation. This was a case of exploiting model outputs, not of stealing the weights themselves.

Comparing weight leaks, distillation, and extraction of data used for fine-tuning, in terms of what is taken and the conditions needed for the attack, gives the following.

Type What is taken Access/conditions required What the attacker can do
Weight leak Parameters of the trained model (RAND's definition) Control of inference servers, etc. (the paper's assumption) With GPU infrastructure, may be able to begin serving within days without additional training (CSA's view)
Distillation Outputs obtained via API (in Anthropic's disclosed case, about 24,000 accounts and over 16 million exchanges) No intrusion into infrastructure needed (CSA) Imitate the original model's capabilities and behavior, but cannot reproduce the weights themselves
Training data extraction Query data used for fine-tuning The creator of a backdoored foundation model has black-box access to the fine-tuned model (ICLR paper) Fully extracted up to 76.3% of 5,000 records under practical experimental conditions (same paper)

The three attacks thus differ in both the information taken and the access required.

Distillation trains a new model using another model's responses as examples, but does not obtain the original weights. Training data extraction targets the input data used in fine-tuning. A weight leak, by contrast, puts the trained model itself, which works as long as it has a suitable runtime environment, into a third party's hands.

Congress demands answers from AI firms: nine companies in April, five in October

On April 29, 2026, Senator Banks and Senate Judiciary Committee Chairman Grassley sent letters to the CEOs of nine companies, OpenAI, Anthropic, Google, xAI, Meta, Microsoft, Amazon, Safe Superintelligence, and Thinking Machines Lab, asking for written responses by May 26.

The questions extended to employee vetting and insider-threat detection. On model weights, they asked whether the companies believe they can sufficiently prevent theft by Chinese threat actors, and how many employees hold privileged access to weights and related sensitive assets and how that number has changed.

They also asked whether the companies plan to notify the U.S. government if they detect or suspect exfiltration by China-linked actors, and if their own AI models or agents attempt exfiltration.

Meanwhile, according to Reuters, Khanna's letters to the five CEOs ask for explanations of known cases aimed at unauthorized access to or theft of weights, as well as the cybersecurity measures in place to prevent such attacks.

Khanna is the ranking Democrat on the House Select Committee on China. Reuters said it did not immediately receive comment from the companies. The congressman warned that Chinese theft of model weights could undermine U.S. leadership in AI "with the stroke of a keyboard."

Organizing the research, company announcements, and congressional actions so far in chronological order gives the following.

Date Event Party
2024 Published "Securing AI Model Weights," outlining 38 attack vectors and five security levels RAND
November 4, 2025 Posted the first version of a research paper on detecting weight exfiltration hidden in inference responses to arXiv Roy Rinberg et al.
February 23, 2026 Disclosed distillation of Claude by three Chinese companies Anthropic
April 29, 2026 Sent letters to nine AI company CEOs asking about weight protection (response deadline May 26) Sen. Banks, Sen. Grassley
May 17, 2026 Published a research note on threat models for AI developers CSA
October 1, 2026 Reported that Khanna asked five CEOs about attempts to steal weights and defenses Reuters

All of these are written inquiries from lawmakers, and no new law or regulation has been decided. Also, the text of Khanna's letter could not be confirmed, so its contents are limited to what Reuters reported.

AD

Hiding weights in AI responses to smuggle them out: the detection method a paper proposes

A five-person research team including Roy Rinberg of Harvard University and MATS posted the first version of the paper "Verifying LLM Inference to Prevent Model Weight Exfiltration" to arXiv on November 4, 2025, and revised it to a third version on March 12, 2026. Co-author Keri Warr is affiliated with Anthropic.

The research aims to raise protection from "SL3" to "SL4" among RAND's five levels of model protection. SL3 is the level expected to stop attacks by cybercrime groups and insiders, while SL4 is the level said to defend against most routine operations by organizations with advanced cyberattack capabilities.

The team assumes an attack in which an attacker who has taken over an inference server mixes weight data into the ordinary responses returned to users and sends it outside.

This method of concealing information is called "steganography." It is a technique that embeds one piece of information within another, carrying secret data in a form that looks like ordinary communication.

The paper points out that data centers operating large-scale inference services must send and receive hundreds of GB of data per day in normal operations alone, and that attackers could use this traffic as cover. The basis for this traffic volume is cited from RAND's report.

However, this is an attack scenario set by the research team; the paper does not report a case in which model weights were actually stolen this way. Reuters also reports that few cases of weights themselves being stolen have been publicly confirmed.

The paper classifies weight exfiltration into three types by time required: fast, completed before humans can respond; intermediate, taking days to weeks; and slow, leaking little by little over months to years.

Sending a large amount of data in a short time is easy to detect. On the other hand, leaking extremely small amounts at a time may take too long even for the attacker. With a detection system in place, attackers must choose between risking detection by sending large volumes or slowing exfiltration to avoid it. The research team focused on how far the latter kind of attack could be suppressed.

At the core of the proposed detection method is re-running and cross-checking inference results.

First, the inputs to the inference server, the generated outputs, and the random-number seed values are recorded. Then a verification server isolated from the outside selects some of the computations and re-runs inference under the same conditions.

Even with the same input, an LLM's output can vary slightly because of differences in the random numbers used for sampling and in floating-point arithmetic. So the verification server does not simply check for an exact match; it estimates whether each generated token falls within the range of variation that could occur with the legitimate model.

The team's key insight is that trying to embed weight information in responses introduces some deviation from the tendencies of output a normal model generates.

The evaluation used four public models from the Llama and Qwen families. It used vLLM as the inference engine and was tested on 500 prompts selected from LMSYS-Chat-1M.

The team reported that for Qwen3-30B-A3B, the amount of information that could be exfiltrated while evading detection was held to under 0.5%. They estimated this corresponds to the time needed for the attack increasing by more than 200 times. For example, an exfiltration that would previously have finished in two days would take more than a year.

That said, this defense has preconditions.

The paper excludes cases of physical theft of servers or storage devices. It also assumes that no side-channel attacks exist and that the verification server has not been compromised. In addition, a footnote explains that it cannot prevent attacks in which the model itself reveals its weights as part of its intended behavior.

The senators' April letters also asked about cases where an AI model or agent itself attempts to exfiltrate weights. This appears close to the kind of attack the new paper does not defend against.

Do companies that fine-tune their own AI models need to protect weights too?

Most of the congressional inquiries and research covered so far mainly target frontier AI companies such as OpenAI and Anthropic, or operators of large-scale inference services.

The inference verification paper deals with inference servers, and CSA's research note centers on AI developers' infrastructure. The recipients of the congressional inquiries are also companies that develop and operate cutting-edge AI models.

However, CSA urges organizations that operate, deploy, or integrate foundation models to treat model weights on par with their most sensitive source code and to put appropriate access controls in place. At the same time, its specific recommendations, such as taking inventory of storage locations, are mainly aimed at organizations that develop, operate, or integrate frontier AI models.

Can the same thinking be applied to ordinary companies? From here, based on public materials, we consider use by Japanese companies.

Take a Japanese company that fine-tunes a public model on its own data and runs it on an internal server. Because the model runs in the company's own environment, that environment stores the weight data.

Even when AI is adopted so as to use it without sending confidential data outside, the model weights themselves become a new information asset that needs protecting.

Those weights reflect the compute resources and tuning work spent on fine-tuning. CSA's point that an attacker with sufficient GPU infrastructure can run stolen weights without additional training would basically apply even when model scale differs.

That said, when the underlying model is publicly available, that part can be obtained by anyone. The value newly handed to a third party through a leak lies mainly in the results of the company's own additional tuning.

On the other hand, the attack scenario assumed in the inference verification paper cannot be applied to ordinary companies as is.

The paper targets the servers of inference providers that return large volumes of responses to outside users. It assumes information is hidden within traffic on the scale of hundreds of GB per day.

By contrast, AI models used only internally likely differ in both traffic volume and communication counterparts. That does not mean internal use is safe, however. Rather than attacks hidden in large volumes of external traffic, it would be necessary to focus on other routes, such as removal by insiders or theft of storage devices.

Backdoors planted in public models could also leak training data

Some research has examined risks specific to companies that fine-tune.

The paper "Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!," accepted to ICLR 2026, showed that if a company fine-tunes a public model with a malicious backdoor, the creator of the original model may be able to extract the data used for the additional training.

All the attacker needs is "black-box access" to the fine-tuned model. This means access solely through exchanges of inputs and outputs, without directly examining the model's internal structure or weights.

In experiments, up to 76.3% of 5,000 query records could be fully extracted under practical conditions. Under ideal conditions favorable to the attacker, the extraction rate reached 94.9%.

What matters is that this is a different attack from weight theft.

It exploits a backdoor planted in a public model to pull out the training data a company added itself, and does not mean that a weight leak necessarily leaks the training data as well.

First, know where weights are stored and who has access

Among the measures CSA proposes for AI developers, one that ordinary companies can relatively easily take on is understanding where model weights are stored.

CSA asks organizations to identify every storage location for mid-training checkpoints and for weights used in production. This covers not just main storage but also backups and experiment-management systems.

It further recommends documenting, for each storage location, who can read, copy, and move files and through which systems, and keeping access logs.

This approach can also be applied by companies operating models they have fine-tuned themselves.

Start by identifying where model weights are stored, and get to a state where you can check the access permissions and operation history for each. Establishing this kind of basic management is likely the first step toward preventing weight leaks.

As for congressional action, we could find no information, within what can be confirmed, on how the companies responded to the senators' April letters. What explanations the companies will give in response to Khanna's latest inquiry is also unknown at this point.

On the technical side, what to watch going forward is whether the inference-verification method proposed in the research paper will be adopted in actual AI service operations.

Under the assumptions that no side-channel attacks exist and that the verification server has not been compromised, it will be important to see whether the research team's result of exfiltration taking more than 200 times longer can be reproduced in real environments.

If that effect is confirmed, it could become considerably harder to steal weights by hiding information in ordinary AI responses. Countermeasures for other routes, such as physical removal and unauthorized access by insiders, would still be essential.

For developers, AI model weights are assets that condense enormous investment, and for companies that do their own additional training, they are important information assets reflecting proprietary know-how. As AI use spreads, how to protect the model itself, not just customer data and internal documents, will become an important issue in corporate security.