Chinese AI Companies Conducting Distillation Campaigns Against U.S. AI Companies [pdf]

michaefe 16 points 14 comments September 08, 2026
media.defense.gov · View on Hacker News

Discussion Highlights (9 comments)

toomuchtodo

Kinda cool and very meta if Chinese AI companies are using distillation to liberate model value from US investment driven AI companies, by then releasing models created with distillation for free. Its circular IP and copyright avoidance and potential model improvement all the way down. Parallels to how China bootstrapped their state manufacturing capacity with Western company joint ventures contributing capital and know how, and then gave them the boot once they no longer needed them.

cyanydeez

O no! China is plundering the very plunder we worked very hard to plunder from everyone on earth.

blackhaz3

Time to run the same playbook back to their frontier labs.

ggm

I wouldn't want to over state it, but from the point of view of a nation state excluded from a class of technology, acquiring it by means which look unusual and probably are is not inherently "wrong" -it probably breaks at least one other nations laws and attracts sanctions but you won't get any party from the USSR or the PRC or any other trade sanctioned economy past and present to agree "they did a wrong thing" because firstly states are essentially amoral and do things in the interest of the state good or bad, right or wrong, they are judged as fruitful or unfruitful and secondly although it's classic "whataboutism" in many ways, "you made me do it" is very strong. "You only leapfrogged us by stealing our tech and copying it" is a very weak line of reasoning. "You did not contribute to the embedded R&D costs to get to this point" is dickering about the price attributed to all three of George Bernard Shaw, Oscar Wilde and Winston Churchill.

mrDmrTmrJ

Can someone walk me through exactly how powerful distillation is? I've never seen a good blog post or metric establishing what % of a new model's "intelligence" can be captured. I'm not disputing the document's claim. Widespread effort at distillation is clear evidence that it is highly effective. I guess I just don't have an intuitive or technical sense for how powerful distillation is.

fakedang

Saw this firsthand when Deepseek's thinking model referenced Claude and Open AI guidelines in separate chats as part of the thinking process notes.

_aavaa_

They’re trying to kidnap what we have rightfully stolen.

segmondy

Distillation is so easy, so why is Google and Meta playing catch up with worse models than many Chinese labs? I mean, they could distill from the Chinese modles which are free? Why are the American labs models not top tier? Arcee, Laguna, Inkling? Why has Cohere or Mistral fallen behind? Our hubris would be the worse of us if we think the Chinese labs are keeping up because of their distillation usages.

leobg

Giving something a sinister name doesn’t change what it is. Didnt Altman say a few years ago that the next frontier for training would be synthetic data? It’s what everyone turned to (Tiny Stories, etc.). Using existing LLMs to generate more training data for new ones is the obvious step. LLMs are offered as document generators. This is what people do. They use them to generate documents. And then they train on those documents. The labs do it themselves. They didn’t ask when they trained on the internet. And now the Chinese - and researchers, and hobbyist, and businesses - dont ask when they train on LLM outputs. Especially when they paid for them. What they call “distillation” is the very knowledge flywheel that we want in society. You buy a book, you may learn from it, and you may write a better one. The author got paid when yiu bought it. Society gets paid when the ideas spread and lead to the creation of more books, products and services, all of which increase the choices that everyone has.

Semantic search powered by Rivestack pgvector
5,917 stories · 53,755 chunks indexed