Article

    Cyber News / Article / Anthropic Says Seven China-Based AI Labs Ran Industrial-Scale Claude Distillation Attacks

    Anthropic Says Seven China-Based AI Labs Ran Industrial-Scale Claude Distillation Attacks
    in
    [email protected] (The Hacker News)-about 4 hours ago

    Anthropic Says Seven China-Based AI Labs Ran Industrial-Scale Claude Distillation Attacks

    Anthropic on Thursday said it identified and disrupted industrial-scale illicit distillation attacks against Claude from seven labs based in China, including Alibaba, Moonshot, DeepSeek, Z.ai (aka Zhipu), and MiniMax.

    Knowledge distillation by itself is alegitimate training method. It refers to amachine learning techniquewhere a large, powerful AI model assumes the role of a "teacher" to train a smaller, less-capable or faster "student" model to copy its capabilities.

    Illicit distillation, on the other hand, is an industrial-scale campaign that covertly extracts a model's capabilities and replicates them in another model without authorization, typically by making use of networks of fake accounts created with stolen credit cards, login credentials, and API keys.

    Frontier AI labs in the West, including those fromGoogleandOpenAI, have repeatedly called out distillation attacks aimed at their models. Anthropic said it has observed unauthorized labs employing "increasingly sophisticated methods" to get around defenses and harvest its capabilities, such as agentic capabilities and tool use, coding and data analysis, and logical reasoning, through prompt manipulation tricks.

    "DeepSeek, Xiaomi, and Moonshot fed conversations between their own models and users into Claude," Anthropicsaid. "These labs then used Claude’s responses as training data with which to distill Claude's capabilities. Some of these exchanges included sensitive information, including from individual users, major multinational companies, and state-affiliated actors."

    The AI company said these labs generally gain access to its models by routing requests through proxy services, also referred to as transfer or relay stations, which create thousands of new accounts under fictitious identities, fake or stolen credit cards, andillegally harvested API keysthat belong to legitimate companies or individuals.

    According to Anthropic, unauthorized AI labs also acquire transcripts of user exchanges with U.S. frontier models by purchasing them off third-party resellers, who are the operators of proxy services that save such conversations without the users' knowledge or consent.

    "In other cases, unauthorized labs rerouted requests from their users to Claude -- without the knowledge or permission of those users -- to harvest exchanges between users and Claude for training," Anthropic pointed out.

    Since February 2026, the AI company said it has detected six illicit distillation campaigns that were conducted by China-based AI labs to advance their own models -

    "The proliferation of proxy services to circumvent Anthropic access restrictions has created a secondary market through which labs can purchase or otherwise acquire harvested exchanges between users and Claude," Anthropic said. "Some proxy networks both provide Claude access to users in unsupported regions, and also save exchanges in order to sell them to other labs."

    To counter illicit distillation, the company said it bans reseller accounts or accounts operating from unsupported regions like China, Iran, and Russia when users fail to verify their identity. To make it harder for unauthorized labs to distill Claude's capabilities, the model has been updated to summarize its internal reasoning before responding, thereby making stolen transcripts less useful for follow-on training.

    "And with Fable 5.1 we introduced preserved thinking, which stops new API accounts from altering the system prompt, tools, or messages that precede Claude's reasoning in multi-turn conversations," the company added. "That reasoning is encrypted, but editing the context before it is a common technique attackers use to make Claude reveal it."

    The development comes as Anthropicsaidit took down a number of accounts that tried to use its models to surveil their citizens and to research diseases in ways that could support biological weapons. Earlier this week, U.S. cybersecurity and intelligence agenciesaccusedChina-based artificial intelligence (AI) companies of conducting "systematic extraction" of proprietary functionalities and capabilities of American frontier models through distillation attacks.

    Original source

    Anthropic Says Seven China-Based AI Labs Ran… | CVE-DB