Anthropic vs China: Why the Distillation Fight Is Real and Hypocritical at Once
Anthropic accused three Chinese labs of distilling Claude through tens of thousands of fake accounts and millions of exchanges. The concern is legitimate, but distillation is also how Anthropic builds its own smaller models, and Claude's training data was scraped from books and a site it was later sued over.

A distillation fight between Anthropic and three Chinese labs made headlines for its scale, roughly 24,000 fake accounts and around 16 million exchanges routed through proxy networks. It is a genuinely interesting story about how AI models get copied, and it is also a strangely useful lens for a small business owner who will never train a model in their life. To show why, I am going to walk one business, a two-dentist practice deciding whether to trust a cheap AI chatbot with patient interactions, all the way through the decision, using the distillation news as the framework that keeps the practice out of trouble.
Chapter one: the practice sees a tempting price
The practice starts where most small businesses start, with a number. The office manager finds an AI chatbot that promises to answer patient questions and handle scheduling for a fraction of what the name-brand options cost. The name-brand tool runs, say, 300 dollars a month. This one is 40. On a small practice's budget, that gap is the entire conversation, and the temptation is to sign up on the spot.
Before signing, though, the owner does one thing that turns out to matter, asks where a tool that cheap actually comes from. That question is the whole case, because the answer to it is exactly what the distillation story explains. Many bargain AI tools are not the real thing. They are distilled or unofficial copies of a bigger model, and understanding what that means is the difference between a smart purchase and a liability sitting between the practice and its patients.

Chapter two: understanding what distillation actually is
So the owner learns how distillation works, because you cannot evaluate a copy without understanding how copies are made. Distillation is when a small student model repeatedly prompts a large teacher model and learns from the answers. Training a frontier model from scratch is brutal, it runs around 90 to 100 days, massive GPU clusters, and hundreds of millions of dollars. Distillation skips almost all of that. The student prompts the teacher hundreds of thousands of times, captures the responses along with the visible reasoning that shows how the teacher thinks, and trains on those prompt-response pairs until the smaller model thinks like the bigger one.
The important thing the owner takes from this is that distillation is not exotic or shady on its own. It is exactly how the major labs produce their own cheaper, faster models, turning a flagship into smaller versions the same way. So the technique itself is normal. What made the news a scandal is a different thing entirely, the accusation that the Chinese labs distilled a model that was not theirs, in a place they technically should not have had access to it. The reported scale is what gives the numbers their punch, with the three labs cited at roughly 0.15 million, 3.4 million, and 13 million exchanges respectively, the largest even pivoting much of its traffic to a new model within a day of its release. For the practice, the takeaway is simpler than the geopolitics, cheap AI tools are often built by copying a bigger model, and a copy is not automatically as safe as the original.

Chapter three: the guardrail question
Armed with that, the owner asks the first hard question about the 40-dollar chatbot, does it keep the safety training that made the original model refuse dangerous requests? This is where distillation gets genuinely risky for a business. Distilled and unofficial copies often lose the safety guardrails the original had, because that safety training does not always survive the copying process. And it compounds, since a model distilled again from an already-stripped copy can end up further from the safe original with each generation down the chain.
For a dental practice, that is not an abstract worry. A chatbot talking to patients might be asked something that edges toward medical advice, and a properly guardrailed model refuses to answer risky medical questions and defers to a professional. A stripped-down copy might just answer, confidently and wrongly, with the practice's name attached to it. The owner realizes the price gap suddenly looks different. The 260 dollars a month saved means very little against one bad automated answer to a patient about medication or symptoms. So the guardrail question becomes a hard requirement, not a nice-to-have. If the vendor cannot clearly show the tool keeps its safety behavior intact, the tool is out, regardless of price.
Chapter four: separating what is allowed from what is legal
The next chapter is subtler, and the distillation news frames it perfectly. The whole industry runs on a take-first, ask-permission-never foundation. The very lab raising the alarm trains its models on data scraped without permission, then objects when someone uses its outputs without permission. The receipts are public, a 1.5 billion-dollar author copyright settlement, roughly 3,000 dollars per book across an estimated 500,000 books, plus a major site suing over data used without payment, and similar accusations against other big labs. The owner is not trying to judge who is right. The owner is extracting the practical lesson buried in it.
That lesson is to treat two questions as separate boxes that both must be ticked. There is what a tool is allowed to do under the terms of service, and there is what is actually legal, and they are not the same thing. The copyright authorities have said AI-generated content is not copyrightable without human authorship, and most terms give output ownership to the user, which is why distilling outputs is a terms-of-service breach rather than a clear crime. Mapped onto the practice, this becomes a concrete vendor-contract question, is the chatbot's handling of patient data actually compliant with health-privacy rules, or just cheap? Compliant and cheap are two separate boxes. The owner decides both must be ticked before the tool ever touches a patient record.
Chapter five: the decision, and the checklist it produced
Here is where the walkthrough pays off. The practice does not just make one yes-or-no call on one chatbot. It walks away with a repeatable habit for evaluating any AI tool, which is worth far more than the single decision. Run the 40-dollar tool through the framework, where does it come from, is it likely a stripped copy, does it keep its safety guardrails, and does it clear both the terms-of-service and the legal-compliance boxes. If it fails the guardrail question or the compliance question, no price makes it worth it. If it passes both, the low price is a genuine win rather than a trap.
In this case, say the practice asks the vendor those questions and gets vague answers on data handling and no clear statement on safety behavior. That vagueness is itself the answer. The practice passes on the bargain tool and either pays for a properly guardrailed option or waits for a vendor that can answer plainly. The 260 dollars a month it might have saved is set against the risk it avoided, and on that scale the math is not close. The framework turned a tempting price into a clear, defensible decision.
The same diligence extends to every other tool the practice touches. The AI writing its patient newsletters and the assistants drafting its Facebook and Instagram ad campaigns deserve the same two-box test before they represent the practice in public, and any tool that plugs into the practice's CRM and website stack, where patient records and follow-up actually live, has to clear the compliance box without exception. The distillation story is really a story about asking better questions before you trust a system, and those questions apply to every piece of software that ever speaks for the business.
Chapter six: the framework becomes the practice's default
The single decision on the chatbot is only the beginning of the story, because the real prize the practice walked away with is a habit it now applies to everything. Once the owner had the four questions, where does the tool come from, is it likely a stripped copy, does it keep its safety guardrails, and does it clear both the terms and the legal-compliance boxes, those questions became the practice's default screen for any new software that touches a patient or speaks for the business.
Consider how differently the practice now moves. When a vendor pitches an AI scheduling assistant, the front-desk staff no longer just compare features and price. They ask the origin question, and if the answer is fuzzy, that fuzziness is treated as a red flag rather than a detail to sort out later. When a marketing tool offers to auto-generate patient education content, the practice checks the guardrail question, because a tool that can confidently produce wrong health information is a liability no matter how nice the content looks. The framework turned a vague sense of caution into a concrete, repeatable checklist, and a checklist is something you can hand to a staff member and trust them to run without you in the room.
There is a subtler shift too, in how the practice reads industry news. Before, a headline like a lab accusing rivals of stealing its model was just noise, a distant fight between big companies that had nothing to do with a dental office. Now the owner reads that same headline for the principle underneath it, and the principle, that the whole industry takes first and asks permission never, and that cheap copies often shed the protections of the originals, is directly useful. The news stopped being entertainment and became market intelligence, the kind that informs a real buying decision. That is a genuine upgrade in how a small business relates to a fast-moving field it cannot control.
And the framework scales to the practice's own obligations, not just its vendors. The two-box test, permission versus legality, is exactly the lens the practice needs for its own compliance posture, because a dental office lives under real privacy rules and cannot afford to treat any tool as compliant just because a salesperson said so. By internalizing that terms-of-service permission and actual legal compliance are separate boxes, the practice built a healthy reflex, verify both before anything touches a patient record, and never assume that cheap and allowed means safe and legal. That reflex protects the practice long after the specific chatbot decision is forgotten.
The distillation drama, in other words, gave a two-dentist office something far more durable than a verdict on one product. It gave the practice a way of thinking about AI tools that will still be useful when the current tools, and the current controversy, are long gone. The specific names will change. The habit of asking where a tool came from and whether it is safe and compliant before trusting it will not, and that habit is the real return on an afternoon spent understanding a news story that looked, at first, like it had nothing to do with running a dental practice.
What the practice actually learned
None of this required the owner to become a lawyer or an AI researcher. It required a framework and the willingness to ask a few pointed questions before signing up. Judge the tool by where it comes from and how it was likely built. Learn enough about distillation to recognize when a cheap tool is a stripped-down copy of something bigger. Confirm the safety guardrails are intact before letting it talk to patients. And treat terms-of-service permission and actual legality as two separate boxes that both have to be checked.
You can absolutely run this diligence yourself, and the practice in this walkthrough did exactly that with nothing more than the right questions. If you would rather have someone vet your AI tools, check the compliance and safety posture, and pick the right ones for your specific business, that is precisely the kind of review an expert can run before anything ever goes live in front of a patient, a client, or a customer.
That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.
Book your call →
