
Today I’m going to explain what AI distillation is, why it frightens the Americans so much (and not only them), and why the Chinese are using it hand over fist to produce AI that is fast, cheap and accurate by exploiting the big models. But I’m also going to explain how much grubby hypocrisy there is behind Anthropic’s tears, and why this whole thing ought to make you angry. Very angry. Some background: as you know, I teach the security of artificial intelligence systems at university, and I always tell my students that in the digital world words don’t describe facts, they frame them in order to define them. This is narrative constructivism, the foundation of an entire career and of my companies, the idea that “reality is a socially negotiated object”, and this is a textbook case of it.
In September 2025, Anthropic, the lab that makes Claude and that presents itself to the world as the “good guys'” artificial intelligence company, the one obsessed with safety and ethical alignment, agreed to pay 1.5 billion dollars to settle a class action brought by American authors: it had downloaded something like seven million books from pirate archives to feed Claude the prose of humanity. A few months later, in February 2026, the very same Anthropic pointed the finger at three Chinese labs, accusing them, with a press release and with numbers, of having stolen its model. The verb used by the press and by the company is always the same: theft. The accusation, in itself, matters little; what matters is that the same word, in the same year, describes two almost identical acts and judges them in opposite ways.
To understand why, it’s worth starting from the technical word that sits at the heart of all this, a word that until two years ago was known only to insiders and that today has become a rhetorical weapon: distillation.
What it means to “distil” a model
Distillation, in machine learning, began as something noble and even remarkably elegant. It was formalised in 2015 by three Google researchers, Geoffrey Hinton (you may remember him by the nickname “the godfather of AI”), Oriol Vinyals and Jeff Dean, in a paper with an almost poetic title, Distilling the Knowledge in a Neural Network. The basic idea is almost disarmingly simple: you have an enormous model, hugely expensive and slow, but very, very good; you want a small model, cheap and fast, that is almost as good but costs less to run. The classic problem of having your cake and your LLM eating it too :). So you take the big model, call it the “teacher”, make it produce a mountain of answers, and you train the small model, the “student”, to imitate something subtler than the mere right-or-wrong verdict: the shades of probability with which the teacher distributes its confidence across all the possible answers. Hinton called this, with a lovely image I’ve always found very poetic, “dark knowledge”: true competence lies not in which the answer is, but in how much the teacher prefers it to the alternatives, and it is in that gradation that the talent is hidden. The student, by learning to imitate those nuances, inherits part of the teacher’s talent at a fraction of the cost.
So far this is a compression technique, stuff for spotty nerds (long live stereotypes!), and within your own four walls it is perfectly legitimate: if teacher and student belong to the same lab, no problem at all. The problem arises when one company’s student goes off, in secret, to take lessons from another company’s teacher. In its 2026 version, in fact, the distillation everyone is talking about is a different thing: you don’t have access to the rival model’s internal mechanisms, but you do have access to its mouth, that is, to its paid public interface. You can “use” it, ask it a great many questions, and exploit it as a teacher… So you open thousands of accounts, you bombard the competitor’s model with questions, you collect millions of its answers, and with that haul you train your own. This is exactly what Anthropic accuses DeepSeek, Moonshot AI and MiniMax of doing: according to the company, behind the accusation lie 24,000 fraudulent accounts and 16 million exchanges with Claude extracted to feed Chinese models, a campaign Anthropic described as being “on an industrial scale”. A few months later, in June 2026, came the twin accusation against Alibaba’s Qwen lab: according to reconstructions by Inc. we’re talking about roughly 25,000 fake accounts and 28.8 million exchanges between April and June, aimed at Claude’s code-writing and reasoning capabilities (this second figure, note well, currently rests on a single chain of sources, and should be taken with the caution it deserves).
Technically, there is a violation: Anthropic’s terms of service, like those of OpenAI, expressly forbid using the model’s outputs to train a competing model. There’s no arguing about this, it’s a contract and it was broken. There is, however, one detail that’s worth the whole story: in the very press release in which it accuses the Chinese, Anthropic admits that distillation is “a widely used and legitimate training method”, only to then argue that the Chinese labs used it “for illicit purposes”. The technique, then, is the same and, by the accuser’s own admission, even legitimate: what changes is merely who wields it. Let’s set this point aside and dwell on the substance of that act, contract aside. To distil a model via its interface means, at bottom, learning by watching what it produces. The student doesn’t copy the teacher’s code, doesn’t steal the network’s weights: it observes its outputs and absorbs its style, its skills, its ways. Keep this definition in mind, because it’ll come back in handy in a moment in an almost embarrassing way.
The theft that comes first
Where, after all, does the prowess of Claude, the robbed teacher, come from? It comes from a far larger and far older distillation, performed on the one teacher no one has ever paid: the entire written output of the human species. To train its models, Anthropic drew, among other sources, on roughly five million books downloaded from Library Genesis and another two million or so from the Pirate Library Mirror, two of the planet’s best-known archives of pirated texts. We know this because it was established by a US federal court in the case Bartz v. Anthropic, settled with that 1.5-billion-dollar agreement announced in September 2025: according to the Authors Guild, the compensation covers roughly 500,000 titles, at around three thousand dollars a book, and it is the highest sum ever seen in a copyright case.
The truly instructive part is the reasoning of Judge William Alsup, filed in June 2025. Alsup drew a clear line: training a model on legally acquired books is, in his words, “quintessentially transformative”, and therefore falls within what American law calls fair use; but downloading the pirated copies of those books is another matter, and is plain and simple copyright infringement. Translation: the problem, for American law, was not learning from the books, it was taking them without paying for them. Anthropic raked in the world’s written knowledge as if it were river water, and found the most convenient rivers in the pirate archives.
The person who described this act with the utmost precision, years before Claude even existed, was Shoshana Zuboff in her The Age of Surveillance Capitalism (2019): digital capitalism, she wrote with an almost unsettling prescience, and with the perfect words to describe the present phenomenon (you can see why I love her, can’t you?), is founded on the “unilateral claim to private human experience as free raw material”. Her analysis concerned behavioural data, clicks and scrolls, but the formula fits language models perfectly: the entire sum of human knowledge treated as free raw material, to be extracted, processed and resold in the form of predictive capacity. With one weighty difference: when Anthropic distils humanity, the court speaks of “transformative” use; when a Chinese firm distils Anthropic, the company speaks of theft. Convenient, isn’t it?
Who decides what is theft? Is property theft?
Here lies the heart of the matter, and it isn’t a hypocritical heart, it’s a political one. In 1840 Pierre-Joseph Proudhon wrote one of the most misunderstood phrases in economic thought, “la propriété, c’est le vol” (“property is theft”). He wasn’t inviting people to steal, mind you! He was observing that every piece of property, at its origin, is an act of appropriation of something that previously belonged to no one, and that only afterwards, with time and with the law, does it become “natural” and respectable. Proudhon’s point is not that stealing is fine; it’s that the difference between legitimate appropriation and theft lies not in the act, it lies in who has the power to name it.
And here words become weapons. Try lining up the vocabulary: when the established company extracts other people’s knowledge, the act is called fair use, raw material, innovation, progress, training. When the challenger extracts the established company’s knowledge, the very same act is called theft, parasitism, and in the Chinese case even a threat to national security: in a memo to the US Congress, OpenAI accused DeepSeek of “free-riding on the capabilities developed by OpenAI and other American frontier labs”. Free-riding. The same logical operation, learning by observing the outputs of a more capable system, changes its name depending on who performs it and against whom.
The political scientist Steven Lukes, in Power: A Radical View (1974), described three dimensions of power. The first is winning open conflicts; the second is deciding what is up for discussion and what stays off the table; but the third, the third face of power, is the deepest: it is the capacity to <u>shape perceptions upstream</u>, so that certain things appear obvious, natural, beyond dispute, and no conflict arises at all. The greatest power is not winning the lawsuit over theft, it’s ensuring that your taking is never called theft by anyone. It is exactly the condition that Antonio Gramsci, in the Prison Notebooks, called hegemony: domination rests not on force, it rests on the ability to turn one’s own worldview into everyone’s common sense, including that of the dominated. “Lawful extraction for me, theft for you” is not a cynical exception to the system of rules: it is the system of rules, written by whoever got there first.
And Michel Foucault completes the picture: in Discipline and Punish (1975) he showed how every regime of power produces its own regime of truth, because knowledge and power hold each other up. The regime that today decides that training on seven million stolen books is “transformative” while querying a public interface is a crime is not discovering a pre-existing truth about theft: it is manufacturing it, and the one who manufactures it is the one who controls the labs, the patents and the press releases. It is, quite literally, the right to be robbed only by me.
It’s not hypocrisy, it’s an enclosure
The easy reading, the one that has already done the rounds online (Cybernews headlined, without mincing words, “Hypocrisy, anyone?”), stops at the word “hypocrisy”, but it’s too comfortable a reading, and a touch lazy. Hypocrisy is an individual moral failing, and identifying a moral failing is the quickest way to fail to understand a structural mechanism. What is happening has a more precise name and a longer history, told by the legal scholar James Boyle in a 2003 essay, The Second Enclosure Movement and the Construction of the Public Domain. The first enclosure, the enclosures, is the one that between the sixteenth and nineteenth centuries in England turned the common pastures, where anyone could bring their sheep, into private property fenced off by hedges and walls. Boyle argued that a second one was under way, this time over intangible goods: knowledge, ideas, culture, fenced off by intellectual property rights and withdrawn from that open territory which is the public domain. It’s the same thesis that Lawrence Lessig advances in Free Culture (2004): the pioneers build their fortunes by helping themselves liberally to the common heritage, and once they reach the top they shut the gate behind them, so that no one can do to them what they did to everyone.
The distillation saga is the second enclosure movement applied to artificial intelligence in its purest form: the commons, here, is the sum of human knowledge: books, articles, code, conversations, the enormous reservoir of intellectual labour accumulated by the species. Whoever arrives first fences it off just once, turns it into a private model, and from that moment defends the fence as legitimate property against anyone who tries to extract from it what it extracted from us. The decisive move is not the initial taking, it’s the conversion: turning an appropriation of the commons into a proprietary asset, and then shifting the conversation. And this is where Lukes’s second face of power, the one of agenda-setting, comes back in handy. The chronology speaks volumes. In January 2025, in his essay On DeepSeek and Export Controls, Dario Amodei, Anthropic’s founder, wrote in black and white “in this essay I take no position on the rumours of distillation from Western models”: distillation, back then, was not the topic (convenient, eh!). In September 2025 comes the 1.5-billion settlement over pirated books. And as if by magic, only in February 2026, four months after paying for its own taking, Anthropic makes other people’s distillation its public battle. Good heavens, how strangely chance plays out, doesn’t it?
Meanwhile that same essay had already shifted the table: the urgent question becomes “how do we stop China from getting the chips to train its models”, and not “by what right, exactly, did we take everyone’s books without asking permission”. One question drives out the other. The downstream theft, the Chinese one, fills the pages; the upstream theft, the founding one, slips off the agenda.
And it does so, mark you, with the tools Lessig had foreseen as far back as 1999 in Code and Other Laws of Cyberspace: laws and courts, of course, but architecture first of all. The gates of the interfaces, the terms of service, the systems for detecting fraudulent accounts, and at the extreme the controls on chip exports, are code and infrastructure that enforce the enclosure better than any ruling. “Minimal government intervention in cyberspace will not mean less regulation”, Lessig wrote: it means only that the regulating, in place of an elected authority, will be done by whoever controls the architecture. And today the architecture of artificial intelligence is controlled by five or six companies.
How to “hold your head high inside this story” (what an awful turn of phrase!)
Beware of falling into the opposite trap, which is equally comfortable: saying that since everyone steals, then no one steals, and that DeepSeek or Alibaba are the honest plunderers of thieves. That’s not how it is, and that’s not the point. Distilling someone else’s model in breach of its terms of service is a genuine wrong, with genuine victims, and the Chinese labs are no Robin Hoods of the algorithm. The point is not to absolve those who imitate: it’s to demand that the rule on theft apply in one direction only until it applies in all directions.
The way to “hold your head high inside it”, as the Americans say, is called symmetry of the rules, and for once it’s a concrete demand, not a wish: if learning by observing the responses of a system is theft when a Chinese lab does it to Claude, then learning from the books of humanity without consent and without compensation is exactly the same theft when Anthropic does it to us: either both are extractions we decide to regulate, with licences, compensation, opt-in and opt-out mechanisms, or neither of them is, and in that case let’s stop invoking national security every time the challenger does to the champion what the champion did to the world. The commons from which everyone has drawn must be governed as a commons, not fenced off by whoever arrived an hour before the others. This, on the policy front, means treating training data as a resource with public rules that are equal for all, and not as a Wild West in which the only law is the speed of the first to plant the flag.
There is, in the end, one sentence that Anthropic could say and never will, because it is at once the most honest and the most uncomfortable. Not “they stole our model”, but <u>”we just got there first”</u>. Everything else, the press releases about Chinese theft, the memos to Congress, the essays about chips, serves to bridge the distance between those two sentences: and that distance, from Proudhon’s enclosed common pasture all the way to the language model, has always been the only thing that separates the pioneer from the thief.
Further reading
- Hinton, G., Vinyals, O. & Dean, J.: “Distilling the Knowledge in a Neural Network” (2015) | https://arxiv.org/abs/1503.02531
- CNBC: “Anthropic accuses DeepSeek, Moonshot and MiniMax of distillation attacks on Claude” (24 February 2026) | https://www.cnbc.com/2026/02/24/anthropic-openai-china-firms-distillation-deepseek.html
- Fortune: “Anthropic claims 3 Chinese companies ripped it off” (24 February 2026) | https://fortune.com/2026/02/24/anthropic-china-deepseek-theft-claude-distillation-copyright-national-security/
- Inc.: “Anthropic Accused Alibaba of a Distillation Attack” (June 2026) | https://www.inc.com/hazel-gandhi/anthropic-accused-alibaba-of-a-distillation-attack-heres-what-that-means-and-why-its-so-dangerous/91365906
- NPR: “Anthropic to pay authors $1.5 billion in settlement” (5 September 2025) | https://www.npr.org/2025/09/05/g-s1-87367/anthropic-authors-settlement-pirated-chatbot-training-material
- PBS NewsHour: “Anthropic to pay authors $1.5B in landmark settlement over pirated chatbot training material” (5 September 2025) | https://www.pbs.org/newshour/nation/anthropic-to-pay-authors-1-5b-in-landmark-settlement-over-pirated-chatbot-training-material
- Authors Guild: “What Authors Need to Know About the Anthropic Settlement” (2025) | https://authorsguild.org/advocacy/artificial-intelligence/what-authors-need-to-know-about-the-anthropic-settlement/
- Dario Amodei: “On DeepSeek and Export Controls” (29 January 2025) | https://darioamodei.com/post/on-deepseek-and-export-controls
- Cybernews: “Hypocrisy, anyone? Anthropic accuses Chinese AI labs of illicit distillation” (2026) | https://cybernews.com/ai-news/anthropic-ai-china-distillation-attack/
- Zuboff, S.: The Age of Surveillance Capitalism (PublicAffairs, 2019)
- Lukes, S.: Power: A Radical View (Macmillan, 1974)
- Gramsci, A.: Prison Notebooks (1929–1935)
- Foucault, M.: Discipline and Punish (Gallimard, 1975)
- Proudhon, P.-J.: What Is Property? (1840)
- Boyle, J.: “The Second Enclosure Movement and the Construction of the Public Domain” (Law and Contemporary Problems, 2003)
- Lessig, L.: Free Culture (Penguin, 2004) and Code and Other Laws of Cyberspace (Basic Books, 1999)
