A striking claim has spread this week alongside the news that US agencies accused six Chinese companies of copying American AI models. Put simply, the story people are telling one another is that you can steal a rival's artificial intelligence just by talking to it. It is a vivid image, and it is worth pulling apart, because the reality is both less magical and more interesting than the headline.
What distillation actually is
The technique at the centre of the row is called knowledge distillation, and it is not a hack or an exploit. A developer sends a large number of prompts to a capable model, collects its answers, and then trains a smaller model on those question-and-answer pairs. The smaller model learns to imitate the larger one's behaviour, the way a student who studies a great teacher's worked examples starts to reproduce the teacher's method. It is a standard, published method taught in machine-learning courses, and labs of every nationality, American ones included, use it routinely to make cheaper, faster versions of their own systems.
What it can and cannot do
Here is the part the theft framing gets wrong. Talking to a model does not give you the model. You cannot read out its weights, the billions of numbers that constitute the trained network, by sending it messages. Those never leave the owner's servers. What you can extract is behaviour, the pattern of how the system responds, and with enough examples you can train a new model that behaves similarly on the tasks you sampled. The copy is an approximation shaped by whatever you thought to ask. It inherits blind spots, and it tends to lag the original rather than match it.
So the honest answer to the question in the headline is a qualified no. You cannot steal an AI by talking to it in the sense of walking away with the thing itself. You can, at sufficient scale, cheaply transfer a good deal of its usefulness into a system you control. Those are very different claims, and collapsing them is what turns a technical practice into a spy thriller.
Then what is the actual dispute?
The genuine argument is not about physics but about permission and scale. Most frontier labs forbid using their outputs to train competing models in their terms of service. The US advisory alleges the copying was industrial in scale and strategic in intent, which is a stronger charge than ordinary research use. China's commerce ministry counters that distillation is a neutral practice everyone relies on. Both things can be true at once. The method is legitimate and widespread, and using it at massive scale against a competitor's paid service may still breach a contract or an emerging norm. That is a legal and ethical question, and it is unsettled, not a matter of settled fact.
The verdict
Treat the shortcut version with caution. Nobody is downloading a rival's brain through a chat box. What is happening is more mundane and, for the companies footing the training bills, more irritating, a cheaper follower learning from an expensive leader's homework. Whether that counts as theft or as the way knowledge has always spread is exactly the fight now playing out between Washington and Beijing. If you want the surrounding context, we have covered China's own moves on export controls and the related worry that models trained too heavily on other models' output can degrade over time.
Sources
- i. thehackernews.com
- ii. www.secureworld.io
- iii. cyberinsider.com
Commentarii · 0