The bottom line: A change to US copyright law could clarify training data usage and ban distillation restrictions, while Chinese manufacturers increasingly pursue open-source releases.
Ben Thompson proposes that the US should legally clarify that data collection for model training falls under Fair Use and prohibit terms of service that ban distillation. This could help US companies compete more effectively with Chinese open-source models.
Thompson argues that the current practice of AI labs is hypocritical: they train their models on data without licensing, but simultaneously prohibit distillation of third-party models in their terms of service — that is, querying them via APIs to extract knowledge.
The proposed legislation would have two components: first, an explicit clarification that collecting training data constitutes Fair Use. Second, a prohibition for US companies to restrict distillation via terms of service. Thompson emphasizes that preventing distillation is technically difficult to enforce, as it amounts to regular API queries.
From a regulatory perspective, the approach would have two effects: it would protect AI labs from liability while simultaneously ensuring that learned knowledge is made available higher up in the innovation chain. This could give US companies, particularly in open-source models, a competitive advantage over Chinese alternatives.
Thompson’s analysis is supported by Alibaba’s recent decision: the company released Qwen 3.8 Max as an open-weights model — a shift from its earlier position this year, when Qwen 3.7 Max was not made openly available. Thompson speculates that this may have been influenced by a speech from Xi Jinping in which he said: “We should seize this rare, historic opportunity to promote open source, openness, collaboration and exchange.”
Source: simonwillison.net · Published July 20, 2026
Lumi AI News — AI-assisted curation pursuant to Article 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.7.3.