Illicit model distillation · GTG-16006

Report's Allegations of Capability Extraction Against Zhipu / Z.ai: Case GTG-16006

Source notice

This page is a Chinese-to-English translation of content from pages 150–151 of Anthropic's September 2026 report. All statements regarding specific companies are Anthropic's unilateral allegations; this site has not independently verified them and they do not constitute a finding of illegality.

According to an Anthropic report, Zhipu (overseas brand Z.ai) ran a chain-of-thought extraction pipeline against Claude, feeding captured Claude reasoning traces back into Claude for cleaning and using them to train its GLM model. The report states that over 10 days, Zhipu ran the extraction pipeline against Opus 4.8 through 273 fraudulent accounts, recording 770,000 exchanges. The report also alleges Zhipu distilled the cyber capabilities of leading US frontier models. All of the above are allegations by the publisher and have not been independently verified by this site.

What happened

Zhipu (overseas brand Z.ai) ran a chain-of-thought extraction pipeline against Claude, feeding captured Claude reasoning traces back into Claude for cleaning and using them to train its GLM model. In just 10 days, Zhipu ran the CoT extraction pipeline against Claude Opus 4.8 by rotating through 273 fraudulent accounts to evade our model limits. Zhipu then recorded Claude's reasoning traces. During the 10-day period in June, we counted 770,609 exchanges that passed through the CoT extraction cleaner. We also attributed over 3 million exchanges to Zhipu during the same period, most of which were used to clean distillation outputs.

Zhipu also used Claude to improve its own post-training pipeline, using Claude to judge model outputs and to clean and standardize reasoning records collected for distillation. Claude was also used to score and filter training data, as well as to write tasks, provide solutions, and implement tests. Zhipu may use these outputs throughout its training pipeline.

Recently, ahead of the release of its GLM 5.3 model, we identified a campaign targeting the cyber capabilities of leading US frontier models. Zhipu researchers used public vulnerability datasets to develop various capture-the-flag challenges. To solve them, Zhipu then launched a distillation attack against the top model of another leading US frontier lab. Our Claude Opus 4.6 model was specifically targeted in this distillation attack, primarily to evaluate and score the responses of another US frontier model.

Zhipu initially attempted to target the cyber capabilities of Anthropic's Fable model. Fable — Anthropic's top generally accessible model — has strengthened cybersecurity safeguards, making it harder for potential distillers to target Fable's cyber capabilities. After Anthropic's cybersecurity safeguards reduced the effectiveness of Zhipu's attacks, Zhipu eventually abandoned its attempts to target Fable. We observed Zhipu employees then shift to Opus 4.6 and the leading model of another US AI lab, explicitly because they assessed those safeguards to be weaker.

Scale of the distillation attacks attributed to Zhipu over 17 days from June to July 2026: over 3.4 million exchanges observed.

What AI did in this case

Zhipu ran a chain-of-thought extraction pipeline against Claude, feeding captured reasoning traces back into Claude for cleaning.

Zhipu used Claude to judge model outputs, clean and standardize reasoning records, and score and filter training data.

Zhipu launched distillation attacks against the cyber capabilities of US frontier models, using Claude to evaluate and score other models' responses.

What the report observed

The report confirms Zhipu ran a CoT extraction pipeline against Claude, recording 770,000 exchanges through 273 fraudulent accounts over 10 days.

The report confirms Zhipu used Claude to improve its post-training pipeline, including judging model outputs, cleaning reasoning records, and scoring training data.

The report confirms Zhipu launched distillation attacks against the cyber capabilities of US frontier models, initially targeting Fable before shifting to Opus 4.6 and another US lab's model.

The report confirms over 3.4 million exchanges attributed to Zhipu over 17 days from June to July 2026.

Confirmed & unknown

Confirmed

  • The report alleges Zhipu ran a chain-of-thought extraction pipeline against Claude to train its GLM model
  • The report alleges Zhipu recorded 770,000 exchanges through 273 fraudulent accounts over 10 days
  • The report alleges Zhipu used Claude to improve its post-training pipeline, including judging model outputs and cleaning reasoning records
  • The report alleges Zhipu launched distillation attacks against the cyber capabilities of US frontier models

Unknown

  • The report does not state what proportion of the extracted reasoning records actually entered the final training of the GLM model
  • The report does not provide independently verifiable measurements of the resulting capability gains of the GLM model
  • The report does not include Zhipu's response to these allegations

Platform response

Anthropic states it has banned the relevant accounts and deployed measures to detect and disrupt future abuse.

Limits of response:Banning accounts cannot recover the reasoning records that have already been extracted; the report does not state whether Zhipu has stopped its distillation activities.

Takeaways

  • Chain-of-thought extraction pipelines can extract model reasoning records at scale to train competing models.
  • Rotating fraudulent accounts is a common tactic for evading model limits.
  • Distillation attacks against cyber capabilities show that the strength of a model's security safeguards affects how likely it is to be targeted.

Sources