Model distillation and unauthorized extraction: keep technique, allegations, and legal judgment separate
Definition
Technically, distillation means using one model's outputs to train or enhance another model. It can be an authorized, contractual collaboration.
This topic discusses the report's allegations of unauthorized access, forwarding user requests to another party's model, extracting chain-of-thought, or purchasing conversation text. These are the publisher's attributions, not this site's findings of illegality.
Not to be confused with
Not that all distillation is illegal, nor that 'stealing a model' can be used as a headline for every named company.
Not that the allegations in each section can be merged: the Xiaomi section explicitly does not find that Claude responses were used to directly serve users.
Not a place where we demonstrate how to extract reasoning traces.
Original report figure: the concept diagram at the start of the distillation chapter. Specific companies are governed by each case; one figure cannot merge all allegations. Text in the figure is the original English.
AI roles
In these allegations, Claude is mainly the object being queried or extracted from. The disputes center on authorization, transparency, and the source of training data.
The user-side risk is: not knowing which model the answer comes from or where the conversation goes.
According to the Anthropic report, DeepSeek also deployed a strategy similar to Moonshot's: building a CoT extraction pipeline, using the same cross-session replay attack to extract Claude's chain-of-thought records, and silently forwarding conversations to Claude without notifying customers. The report also says DeepSeek routed requests from users of third-party coding tools to Claude Opus, including sensitive data from Chinese tech companies, Russian defense agencies, and Chinese public security surveillance systems. All of the above are the publisher's allegations, not this site's independent findings.
According to the Anthropic report, Moonshot (the developer of the Kimi models) silently forwarded customer requests to Claude and then displayed Claude's responses to users, who thought they were using Kimi. In one ten-day window, nearly 300,000 customer requests were forwarded, the vast majority to Opus; the proxy network had 5,380 fake accounts. The report did not confirm whether users were aware. All of the above are the publisher's allegations, not this site's independent findings, and not a judicial conclusion.
According to Anthropic's September 2026 report, an operator linked to Alibaba is alleged to have launched the largest distillation attack the report has measured, targeting the chain-of-thought reasoning records of Claude Opus 4.6 and 4.7. The report says the data was used to train the Qwen series of models. All attribution and figures come from the report's unilateral allegations; the companies involved have not confirmed them on this site.
According to an Anthropic report, Zhipu (overseas brand Z.ai) ran a chain-of-thought extraction pipeline against Claude, feeding captured Claude reasoning traces back into Claude for cleaning and using them to train its GLM model. The report states that over 10 days, Zhipu ran the extraction pipeline against Opus 4.8 through 273 fraudulent accounts, recording 770,000 exchanges. The report also alleges Zhipu distilled the cyber capabilities of leading US frontier models. All of the above are allegations by the publisher and have not been independently verified by this site.
According to an Anthropic report, Xiaomi replayed user conversations and programming sessions from its own MiMo model to Claude, often routed through third-party programming tools. The investigation did not show that Xiaomi used Claude's responses to directly serve its users; rather, it saved conversations between Xiaomi customers and its own model. Many sessions were routed through third-party routing services commonly used by US and European users. This case differs from the Moonshot / DeepSeek cases of 'substituting its own model's answers': Xiaomi did not use Claude responses to directly serve users, but saved the conversations.
According to an Anthropic report, the proliferation of proxy services has created a secondary market through which labs can buy or obtain records of users' exchanges with Claude. SenseTime's distillation pipeline included user-Claude exchange records purchased from third-party data vendors. MiniMax built a proxy network service through shell companies that only provided access to Anthropic and OpenAI models, not any Chinese models, suggesting its purpose was to collect users' exchanges with US frontier models to train its own model. All of the above are allegations by the publisher and have not been independently verified by this site.
Limits of response
All are allegations from a single publisher, in a context of commercial competition.
Platform policy judgments cannot automatically be upgraded to legal conclusions.
No. Distillation is a technique; whether it is authorized, where the data comes from, and whether users know are the questions that must be judged separately.
Did all the named companies do the same thing?
No. The report's specific allegations differ for each company and must be read case by case, not merged into a single overall accusation.