On July 22, 2026, Michael Kratsios, Director of the White House Office of Science and Technology Policy (OSTP), directly named China's Moonshot AI, alleging that the company distilled Anthropic's Fable and used it to develop Kimi K3. He further stated that the company had acquired servers equipped with NVIDIA GB300 chips and had also accessed GB300 in Thailand. In April, a government document had already included warnings about organized distillation centered on China, but this time a specific company, model, and chip were identified. However, Kratsios has not disclosed logs showing that Fable's outputs were incorporated into K3, nor has he disclosed the procurement route for the GB300s.

What this allegation raises is the provenance of a model—something that cannot be judged from benchmark numbers alone. Distillation itself is a widely used development technique. What is being disputed are two separate paths: whether the terms of service and regional restrictions were circumvented to extract a competing model's capabilities at scale, and how access was obtained to export-controlled computing resources.

AD

Two allegations the US government laid out side by side: Fable distillation and GB300

In a post on X, Kratsios stated that the US government understands Moonshot AI distilled Fable in the course of developing K3. According to him, Moonshot AI built an in-house infrastructure for large-scale distillation of US-made models and was able to quickly switch between multiple access methods to avoid detection. "Distillation" here refers to a technique in which the outputs of a strong teacher model are used as training data to transfer capabilities to a different model.

The same post also touched on the chip supply route. Kratsios claimed that Moonshot AI had acquired GB300-equipped servers and had also accessed GB300 chips located in Thailand, likely using them to train AI models. Training and post-training processes that transfer a teacher model's outputs to a different large-scale model require computing resources. In a single post, Kratsios laid out two distinct allegations side by side: how model outputs were obtained, and how GPU access was secured.

However, neither claim has reached the stage of proof at this point. The published post does not show Anthropic-side access logs, indicators remaining in K3's training data, or details of who owned the servers and for how long they were used. The US government's allegation has become more specific, but evidence that third parties could use to independently verify it has not yet been made public.

What separates legitimate distillation from an "attack" is the manner of access

Anthropic itself has explained that distillation is a legitimate technique routinely used across the industry to create smaller, cheaper models. When a company creates a smaller version from its own teacher model, or trains on outputs it has been authorized to use, the technique itself is not the problem. What Anthropic calls a "distillation attack" is the act of bundling fake accounts and proxies to extract specific capabilities at a volume and repetition pattern that differs from normal usage.

On February 23, Anthropic announced that three companies—DeepSeek, Moonshot AI, and MiniMax—had used approximately 24,000 fraudulent accounts to conduct more than 16 million interactions with Claude in total. Of these, more than 3.4 million were attributed to Moonshot AI, which Anthropic said used hundreds of fraudulent accounts and multiple access routes. The capabilities targeted included agentic reasoning and tool use, coding and data analysis, and computer vision.

As grounds for attribution, the company cited correlations in IP addresses, request metadata, and infrastructure-level indicators. Regarding Moonshot AI specifically, Anthropic said the metadata matched publicly available profiles of the company's executives. This explanation provides concrete material that predates the current allegation. On the other hand, the February announcement dealt with activity around Claude in general and Kimi-series models, and is not a report that directly links Fable and K3, which emerged in July. The figure of 3.4 million interactions alone does not support the conclusion that K3 distilled Fable.

AD

The allegation followed the release of the 2.8-trillion-parameter K3

On July 16, Moonshot AI unveiled Kimi K3 as a model with 2.8 trillion parameters, a context length of 1 million tokens, and native vision capabilities for handling images. It adopts a Mixture of Experts architecture that activates 16 out of 896 experts, and the company describes it as "the first open 3-trillion-class model." Since Kimi K2 had a total of 1 trillion parameters, this represents roughly a 2.8x increase in the scale of the open model.

The significance of K3 is not determined by scale alone. Moonshot AI claims that, according to its own evaluations, while K3 falls short of Fable 5 and GPT-5.6 Sol in overall performance, it consistently outperforms other models it was evaluated against. The official API pricing is $3 per million input tokens (non-cached) and $15 per million output tokens. While this pricing allows one to calculate the cost of API usage, comparing total costs including self-hosted operation will require the release of all model weights and measured data.

That said, closeness in benchmarks does not prove the provenance of training data. Moonshot AI's evaluations involved different execution environments depending on the model, and it is noted that fallbacks occurred in some evaluations of Fable 5. The company plans to release the full model weights for K3 by July 27, along with a technical report detailing architecture, training, and evaluation. Even after release, the training data itself may not necessarily be disclosed, but at the very least, outside researchers will be able to closely examine the weights and behavior.

Verifiable evidence needed before sanctions

In an April 23 national science and technology memorandum, the US government indicated its view that foreign entities—particularly those based in China—were distilling US frontier AI at an industrial scale. A subsequent April 29 letter from a congressional committee also cited the roughly 24,000 accounts and more than 16 million interactions, and has begun investigating development tools that incorporate Chinese-made models. The July 22 post connects this policy stance to K3, the latest product to emerge.

Treasury Secretary Scott Bessent said on Fox Business the day before that the US has the ability to impose sanctions if it is confirmed that a foreign model stole technology from a US company. According to AFP, he referenced distillation and claimed that "watermarks" of US-made large language models have been found in numerous Chinese-made models, indicating an intention to investigate within days to weeks. The targets of sanctions, the legal basis, and the certification process have not yet been made public.

The next policy decision will require more than repeated allegations. Time-stamped logs showing access to Fable, attribution information linking the account clusters to Moonshot AI, technical indicators connecting extracted outputs to K3's training process, and records of GB300 ownership and usage—once these are assembled, the allegation becomes verifiable. The first confirmable milestone is July 27. Whether K3's full weights are released as scheduled, how much the subsequent technical report explains about training and evaluation, and whether the US government presents evidence or takes action will determine where the line between competition policy and enforcement ultimately falls.