On July 22, White House science adviser Michael Kratsios did something unusual. He logged onto X and accused a specific Chinese company, Moonshot AI, of stealing a specific American model, Anthropic's Fable, through a specific method: large-scale distillation via API abuse.

A dark-cloaked figure pours sand into a merchant's palm in a crowded marketplace with uneven scales in the background.

This was not a legal filing. It was not a corporate blog post. It was a White House official drawing a line in public. Treasury Secretary Scott Bessent followed within hours, posting on X that sanctions "will be on the table" and that he was considering adding Moonshot AI to a trade blacklist. The Chinese embassy in Washington, through spokesperson Liu Chang, called the allegations "entirely unfounded."

The dispute between AI labs is now a US government operation. The stakes, as Anthropic public policy executive Sarah Heck put it the same day, are national security: "Chinese theft of U.S. models creates serious national security risks for the United States."

Armored figures lower an iron gate over a chasm while robed figures study a map by torchlight on the far side.

The 16 million query pipeline

Moonshot AI unveiled its Kimi K3 model on July 16, 2026: 2.8 trillion parameters, a context window of one million tokens, native multimodal capabilities. Six days later, Kratsios posted his accusation.

A cartographer with a compass stands on a cliff overlooking a stormy coastline, with a ship and a raven nearby.

The speed of the US response signals coordination. But the evidence had been accumulating for months.

In February 2026, Anthropic publicly disclosed that three Chinese firms—DeepSeek, Minimax, and Moonshot AI—had created over 24,000 fraudulent accounts and issued more than 16 million queries to its Claude API. A Cloud Security Alliance research note published March 8 confirmed the scale: more than 16.55 million queries through approximately 24,000 fraudulent accounts, all aimed at extracting Claude's capabilities through knowledge distillation.

There was no breach. The CSA note is explicit: Anthropic found no evidence of unauthorized access to internal systems. The attackers bought legitimate API access through fraudulent account clusters that evaded geographic restrictions. They used the API exactly as designed, at industrial volume, with structured prompts calibrated to transfer specific capabilities.

This is the core of the problem. The attack vector is not a penetration. It is the business model.

How distillation became a weapon

Model distillation is a legitimate technique. A smaller "student" model queries a larger "teacher" model, learns from the responses, and transfers capabilities without touching source code or model weights. Inside a single lab, it is how you make large models efficient.

Across company lines, at scale, it becomes extraction. The student is a competitor. The teacher is an unwitting R&D subsidy.

Kratsios alleged that Moonshot built "a sophisticated internal platform to conduct large-scale distillation against U.S. models" that allowed it "to quickly switch between multiple methods of access to avoid detection." This was not a scrappy effort. It was infrastructure purpose-built for evasion.

The consensus response frames this as a clear-cut case of Chinese IP theft. That framing is accurate but incomplete. It misses a harder question: why was the pipeline so open?

Anthropic sold API access. The fraudulent accounts were a nuisance, not a penetration. The attackers paid for the queries that extracted the capabilities. This does not excuse the theft. But it reveals an architectural vulnerability that no lawsuit can close. The API is the product, and the API is the attack surface. The two are the same thing.

The hardware trail runs through Thailand

Kratsios also stated that Moonshot gained access to restricted Nvidia servers, including GB300 chips, and used servers in Thailand to train its models.

This is the second front. The US restricts advanced chip sales to China. Chinese firms rent or buy access in countries with looser controls. The model is trained on restricted hardware, using capabilities extracted from restricted APIs. The output is a 2.8 trillion-parameter model that competes with the best American labs have built.

The US government is now responding to both fronts simultaneously: the API pipeline and the hardware pipeline. Kratsios' accusation is the public opening of a campaign that has been building since at least February.

What happens next: sanctions, blockers, and a bifurcated ecosystem

The US will impose sanctions on Moonshot AI within 12 months. Bessent's statement was not hypothetical. Treasury has the authority. The political cover is bipartisan. The evidence is public. Adding Moonshot to a trade blacklist is the lowest-cost signal the administration can send. The only question is timing.

Sanctions are the stick. The shield is a fundamental shift in how American AI companies defend their APIs.

Within 18 months, major US labs will deploy real-time distillation detection systems. The current model—sell API access, monitor for abuse, file complaints after the fact—has failed. Sixteen million queries is not a monitoring gap. It is a structural failure.

The next generation of API infrastructure will treat every query as a potential extraction attempt. Rate limiting will become dynamic and behavioral. Query patterns will be fingerprinted. Accounts that exhibit distillation behavior—structured, repetitive, capability-targeted prompting at scale—will be throttled or cut off automatically. This requires inference-time overhead, behavioral modeling, and shared threat intelligence across labs. It is expensive. The alternative is continuing to subsidize the competition's R&D.

Export controls will tighten further. The Thailand server revelation will accelerate efforts to track and restrict Nvidia hardware in third countries. The US will pressure allies to enforce chip controls more aggressively. Chinese firms will respond by investing in domestic chips and cloud proxies. The cost of compute access will rise. The cost of extraction will rise with it.

The second-order consequence is a bifurcated global AI ecosystem. US-aligned labs will share defensive infrastructure and threat intelligence. China-aligned labs will develop alternative compute and more sophisticated evasion techniques. Distillation will remain the primary vector of conflict because it is cheaper than fundamental research and faster than organic development.

The arms race does not end. It moves to the API gateway.

What operators should do now

AI lab executives and policymakers have a narrow window to harden infrastructure before regulation forces their hand.

Invest in API monitoring and dynamic rate-limiting systems now. Collaborate on shared threat intelligence for distillation patterns. Prepare for a regulatory environment where model weights are treated as controlled munitions and API access is subject to export controls.

For Chinese firms, the message is clear: compute access and cloud accounts will face increased scrutiny. The era of buying API access through fraudulent accounts and training on restricted hardware in third countries is closing. The cost of doing business just went up.

The line is drawn

Kratsios' accusation is a line in the sand. The Chinese embassy's denial is predictable. The evidence is harder to dismiss: 16 million queries, 24,000 fraudulent accounts, a 2.8 trillion-parameter model trained on restricted hardware in Thailand.

IP theft at scale is now a national security matter. The US is building the infrastructure to fight it—not in courtrooms, but in API gateways and export control offices. The ledger is public. The response is coming.