Unauthorized model distillation

Model distillation and unauthorized extraction: keep technique, allegations, and legal judgment separate

Definition

Technically, distillation means using one model's outputs to train or enhance another model. It can be an authorized, contractual collaboration.

This topic discusses the report's allegations of unauthorized access, forwarding user requests to another party's model, extracting chain-of-thought, or purchasing conversation text. These are the publisher's attributions, not this site's findings of illegality.

Not to be confused with

Not that all distillation is illegal, nor that 'stealing a model' can be used as a headline for every named company.

Not that the allegations in each section can be merged: the Xiaomi section explicitly does not find that Claude responses were used to directly serve users.

Not a place where we demonstrate how to extract reasoning traces.

Diagram: the service users think they are getting may differ from where the data goes, as described in the report.
Original report figure: the concept diagram at the start of the distillation chapter. Specific companies are governed by each case; one figure cannot merge all allegations. Text in the figure is the original English.

AI roles

In these allegations, Claude is mainly the object being queried or extracted from. The disputes center on authorization, transparency, and the source of training data.

The user-side risk is: not knowing which model the answer comes from or where the conversation goes.

Related cases

Illicit model distillation

Report's allegations about DeepSeek request routing: GTG-16001 case

According to the Anthropic report, DeepSeek also deployed a strategy similar to Moonshot's: building a CoT extraction pipeline, using the same cross-session replay attack to extract Claude's chain-of-thought records, and silently forwarding conversations to Claude without notifying customers. The report also says DeepSeek routed requests from users of third-party coding tools to Claude Opus, including sensitive data from Chinese tech companies, Russian defense agencies, and Chinese public security surveillance systems. All of the above are the publisher's allegations, not this site's independent findings.

Illicit model distillation

Report's allegations about Moonshot request routing: GTG-16002 case

According to the Anthropic report, Moonshot (the developer of the Kimi models) silently forwarded customer requests to Claude and then displayed Claude's responses to users, who thought they were using Kimi. In one ten-day window, nearly 300,000 customer requests were forwarded, the vast majority to Opus; the proxy network had 5,380 fake accounts. The report did not confirm whether users were aware. All of the above are the publisher's allegations, not this site's independent findings, and not a judicial conclusion.

Illicit model distillation

Report's allegations about Alibaba-related distillation activity: GTG-16005 case

According to Anthropic's September 2026 report, an operator linked to Alibaba is alleged to have launched the largest distillation attack the report has measured, targeting the chain-of-thought reasoning records of Claude Opus 4.6 and 4.7. The report says the data was used to train the Qwen series of models. All attribution and figures come from the report's unilateral allegations; the companies involved have not confirmed them on this site.

Illicit model distillation

Report's Allegations of Capability Extraction Against Zhipu / Z.ai: Case GTG-16006

According to an Anthropic report, Zhipu (overseas brand Z.ai) ran a chain-of-thought extraction pipeline against Claude, feeding captured Claude reasoning traces back into Claude for cleaning and using them to train its GLM model. The report states that over 10 days, Zhipu ran the extraction pipeline against Opus 4.8 through 273 fraudulent accounts, recording 770,000 exchanges. The report also alleges Zhipu distilled the cyber capabilities of leading US frontier models. All of the above are allegations by the publisher and have not been independently verified by this site.

Illicit model distillation

Report's Allegations of Training Data Use by Xiaomi: Case GTG-16008

According to an Anthropic report, Xiaomi replayed user conversations and programming sessions from its own MiMo model to Claude, often routed through third-party programming tools. The investigation did not show that Xiaomi used Claude's responses to directly serve its users; rather, it saved conversations between Xiaomi customers and its own model. Many sessions were routed through third-party routing services commonly used by US and European users. This case differs from the Moonshot / DeepSeek cases of 'substituting its own model's answers': Xiaomi did not use Claude responses to directly serve users, but saved the conversations.

Illicit model distillation

Allegations of Third-Party Resale and Training Data Ecosystem: Cases GTG-16012/16003

According to an Anthropic report, the proliferation of proxy services has created a secondary market through which labs can buy or obtain records of users' exchanges with Claude. SenseTime's distillation pipeline included user-Claude exchange records purchased from third-party data vendors. MiniMax built a proxy network service through shell companies that only provided access to Anthropic and OpenAI models, not any Chinese models, suggesting its purpose was to collect users' exchanges with US frontier models to train its own model. All of the above are allegations by the publisher and have not been independently verified by this site.

Limits of response

All are allegations from a single publisher, in a context of commercial competition.

Platform policy judgments cannot automatically be upgraded to legal conclusions.

Related guides

FAQ

Is distillation just stealing a model?

No. Distillation is a technique; whether it is authorized, where the data comes from, and whether users know are the questions that must be judged separately.

Did all the named companies do the same thing?

No. The report's specific allegations differ for each company and must be read case by case, not merged into a single overall accusation.