US agencies accuse Chinese AI firms of distilling US models
TECH

US agencies accuse Chinese AI firms of distilling US models

39+
Signals

Strategic Overview

  • 01.
    The NSA, CISA, and FBI released a joint cybersecurity advisory (AA26-251A) on September 8, 2026, accusing six China-based AI companies of conducting industrial-scale knowledge distillation against U.S. frontier AI models since at least late 2024.
  • 02.
    The advisory names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI as the six companies conducting the campaigns, targeting variants of Anthropic's Claude, OpenAI's GPT, Google's Gemini, and xAI's Grok.
  • 03.
    The agencies describe the campaigns as core to the accused companies' AI development strategy, not merely a supplementary practice, and say the requests were routed through native APIs, remote cloud providers, and third-party aggregators that obscure user metadata, violating providers' terms of use.
  • 04.
    The advisory identifies a gray market of proxy services called 'transfer stations,' reportedly marketed on Chinese platforms Taobao and Xianyu, used to bypass geographic restrictions and evade provider safeguards.
  • 05.
    CISA's Acting Director publicly urged AI companies to safeguard their platforms against these distillation campaigns.
  • 06.
    Separately, ahead of the government advisory, Anthropic reported that DeepSeek, Moonshot, and MiniMax created roughly 24,000 fake accounts and generated over 16 million interactions with Claude to extract capabilities.

Deep Analysis

Inside the advisory: how the alleged extraction pipeline works

On September 8, 2026, the NSA, CISA, and FBI released a joint cybersecurity advisory, AA26-251A, accusing six China-based AI companies of conducting industrial-scale knowledge distillation against US frontier models since at least late 2024 [1]. Knowledge distillation itself is not exotic: a smaller 'student' model is trained to mimic the outputs of a larger 'teacher' model, learning to approximate its reasoning patterns without needing the teacher's original training data or compute budget. What the advisory alleges is different in scale and method. It says the accused firms treated distillation as 'the core - not merely a supplement - of their AI development strategy,' and routed requests through native APIs, remote cloud providers, and third-party aggregators specifically chosen to obscure user metadata and evade the target companies' terms of use [2]. The advisory also identifies a gray market of proxy services called 'transfer stations,' reportedly marketed on Chinese platforms Taobao and Xianyu, that bypass geographic restrictions and safeguards, alongside bulk purchases of premium subscriptions shared across developer teams to keep the cost of extraction low [3].

Named and shamed: who is accused of taking what

The advisory names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI, and says the campaigns extracted billions of tokens across millions of exchanges from variants of Anthropic's Claude, OpenAI's GPT, Google's Gemini, and xAI's Grok [1]. DeepSeek and Moonshot AI are named as the top offenders: DeepSeek is accused of distilling multiple versions of Claude, Gemini, ChatGPT, and Grok to train its R1 model, while Moonshot AI allegedly distilled 18 different US models to build Kimi-K2 and Kimi-K3 [3]. Alibaba is accused of using industrial-scale distillation to improve its Qwen family, and Z.AI allegedly pulled billions of tokens of chain-of-thought reasoning from GPT-5.5 and Claude Opus 4.8 by mid-2026 [4]. Separately, and ahead of the government advisory, Anthropic said it had identified roughly 24,000 fake accounts and more than 16 million interactions with Claude used by DeepSeek, Moonshot, and MiniMax to extract its capabilities [5]. The advisory also takes a swipe at DeepSeek's credibility, questioning its widely publicized $5.6 million training-cost figure on the grounds that it excludes the value of the distilled training data pulled from US frontier models [4].

Theft or standard practice? The framing fight

The advisory's language - 'systematic extraction,' 'industrial-scale' - casts distillation as adversarial IP theft, and CISA's leadership has echoed that framing publicly. But the claim is far from uncontested. Chinese state media has pushed back directly: a Global Times-cited commentator dismissed Anthropic's distillation claims as lacking substance and argued they are 'rooted in tech hegemony anxiety,' framing the accusations as an attempt to protect a fading US monopoly rather than address genuine wrongdoing [6]. Independent academics occupy more cautious middle ground - Erik Cambria, an AI professor at Nanyang Technological University, has noted that 'the boundary between legitimate use and adversarial exploitation is often blurry' [7], since distillation via paid API access is a widely used, technically unremarkable machine-learning technique rather than a hack. That ambiguity matters because the advisory's own evidence describes paying customers using publicly available APIs and cloud services, not intrusion or credential theft - the alleged violation is of terms of use and access-control evasion, not a breach in the traditional cybersecurity sense.

Why Washington is alarmed now, and what comes next

The government advisory did not emerge in isolation. OpenAI first told Axios in January 2025 that it had evidence of distillation attempts by China-based groups targeting its models, shortly after DeepSeek's R1 release drew scrutiny [8]. By February 2026, OpenAI told Congress that DeepSeek had used distillation and obfuscated routers to 'free-ride' on US frontier-lab capabilities, the same month Anthropic separately reported industrial-scale distillation campaigns by DeepSeek, Moonshot AI, and MiniMax against Claude [9]. That reporting prompted a policy debate in Washington by July 2026 over penalties, export controls, and legal safeguards for AI labs [10]. By early September, Anthropic said its fight against distillation had extended into dark-web-adjacent gray markets just before the joint advisory landed [11]. Officials frame the stakes in stark competitive terms: unchecked distillation, they warn, threatens to close the capability gap that has defined US AI leadership [12]. CISA's Acting Director has publicly urged AI companies to 'take immediate steps to safeguard their platforms against knowledge distillation campaigns that threaten to close the gap in advancements made by American companies' [4]. That urgency is bolstered by an independent data point circulating in tech commentary: aggregate model-usage data reportedly showed Chinese open-source model traffic climbing from roughly a fifth to about half of total usage over the past year, while US model usage over the same period slid from about three-quarters to roughly a third - two separate trend lines, not a clean one-for-one swap, that together critics say reflect price competition as much as any distillation-driven capability catch-up, alongside a persistent counter-argument that American labs built their own frontier models by scraping the open web without payment or consent in the first place.

Historical Context

2025-01
OpenAI told Axios it had evidence of distillation attempts by China-based groups targeting its models, shortly after DeepSeek's R1 release drew scrutiny.
2026-02
Anthropic said it identified industrial-scale distillation campaigns by DeepSeek, Moonshot AI, and MiniMax against Claude; OpenAI separately told Congress DeepSeek used distillation and obfuscated routers to 'free-ride' on US frontier lab capabilities.
2026-07
Anthropic and OpenAI warnings about distillation prompted a policy debate in Washington DC over how to respond.
2026-09-03
Anthropic reported its distillation battle with Chinese firms extending to dark-web-adjacent gray markets as concerns intensified ahead of the government advisory.
2026-09-08
The three agencies released joint advisory AA26-251A naming six Chinese AI companies and detailing distillation methods, targeted models, and mitigation recommendations.

Power Map

Key Players
Subject

US agencies accuse Chinese AI firms of distilling US models

NS

NSA, CISA, FBI (US government)

Issued the joint advisory (AA26-251A) accusing the six firms and outlining detection/mitigation recommendations

DE

DeepSeek

Accused top offender; alleged to have used distillation to train its R1 model, undermining its claimed low training cost

MO

Moonshot AI

Accused top offender; alleged to have distilled 18 U.S. models to train Kimi-K2/Kimi-K3

AL

Alibaba

Accused of using industrial-scale distillation to improve its Qwen model family

MI

MiniMax

Accused of targeting Claude, Gemini, and GPT for agentic coding and tool orchestration

ST

StepFun

Accused of targeting Claude and GPT models to improve its products

Z.

Z.AI

Accused of distilling billions of tokens from GPT-5.5 and Claude Opus 4.8 by mid-2026

AN

Anthropic

Target company; separately disclosed distillation campaigns against Claude and is lobbying for penalties and export controls

OP

OpenAI

Target company; earlier told Congress it had evidence of distillation by China-based groups, describing DeepSeek's approach as 'free-riding' on U.S. frontier lab capabilities

GO

Google (Gemini) / xAI (Grok)

Additional target model providers named in the advisory alongside Claude and GPT

CH

Chinese government

Advisory assesses the campaigns 'likely' occurred with Chinese government awareness, though it stops short of alleging direct control

Fact Check

12 cited
  1. [1] Cybersecurity Advisory AA26-251A: China-Based AI Companies' Malicious Knowledge Distillation Against U.S. AI Companies
  2. [2] US claims Chinese AI companies' core AI strategy is distilling American models
  3. [3] US accuses Chinese AI companies of distillation
  4. [4] China: malicious AI knowledge distillation against US companies
  5. [5] Chinese AI companies distilled Claude to improve their models, Anthropic says
  6. [6] Anthropic's 'distillation' claims against Chinese firm lack substance, rooted in tech hegemony anxiety
  7. [7] The AI cold war: US tech companies accuse China's AI firms of stealing billions in research
  8. [8] OpenAI-DeepSeek distillation dispute deepens US-China AI tensions
  9. [9] Anthropic, OpenAI say China AI firms used distillation, including DeepSeek
  10. [10] Anthropic, OpenAI warnings prompt distillation debate in Washington
  11. [11] Anthropic's distillation battle turns to the dark web as China concerns swell
  12. [12] CISA warns of Chinese AI firms' industrial-scale distillation

Source Articles

Top 5

THE SIGNAL.

Analysts

Publicly urged AI companies to take immediate protective action against distillation campaigns to prevent Chinese firms from closing the capability gap with American AI companies.

Nick Andersen, CISA Acting Director
US government official

Says distillation campaigns against its models are intensifying and becoming more sophisticated, framing this as a security threat broader than any single company or country.

Anthropic (corporate statement)
Targeted AI lab

Cautions that the line between legitimate AI research practice and adversarial exploitation is not always clear-cut.

Erik Cambria, professor of artificial intelligence, Nanyang Technological University (Singapore)
Independent academic analyst

Argues Anthropic's distillation claims lack substance and reflect US anxiety about losing technological dominance.

Unnamed Chinese expert cited by Global Times
Chinese state-media-cited commentator
The Crowd

China-based AI companies are illicitly distilling U.S. frontier AI capabilities. Read NSA's new report, co-sealed with @FBI @CISAgov, highlighting AI knowledge distillation, TTPs used, and recommended mitigations: media.defense.gov/2026/Sep/08/2...

@@NSACyber731

China-based AI companies are conducting industrial-scale distillation campaigns to extract proprietary capabilities from U.S. AI models. Our joint advisory with @NSAGov @FBI details the threat and actions to help protect U.S. AI leadership. Read more → go.dhs.gov/56v

@@CISAgov104

US security agencies accuse Chinese AI companies of improperly piggybacking on leading American AI models using a technique known as distillation. Here's what to know

@@business17

U.S. agencies say Chinese AI companies are conducting industrial-scale distillation of U.S. frontier models

@u/Cklly200494
Broadcast
Behind China's rapid AI ascent

Behind China's rapid AI ascent

Why Anthropic is right to fear Chinese AI distillation

Why Anthropic is right to fear Chinese AI distillation

USA Fighting Chinese AI Distillation Attacks - America Losing to China

USA Fighting Chinese AI Distillation Attacks - America Losing to China

US agencies accuse Chinese AI firms of distilling US models — AI News | Agentic Brew