Skip to content

AI Developers Buy Books Published Before 2022 in Bulk for Training Data

In a nutshell: AI companies deliberately purchase older books as training data and destroy them during the digitization process.

AI enterprises acquire large quantities of out-of-print or antiquarian books with publication dates before 2022 to use their contents as AI-free training data. The books are destroyed when scanned.

AI developers are increasingly turning to a strategy to obtain legally secure training data: they are buying books with publication dates before 2022 on a massive scale. Such older printed works are considered free of AI training material and thus enable a clear distinction from data that may have already been influenced by generative AI systems.

The purchasing process often targets out-of-print, antiquarian, or difficult-to-access works that can be obtained at low prices. When the books are scanned, they are typically destroyed afterwards—a practice that is both cost-effective and irreversible, raising questions about resource efficiency.

This approach reflects the regulatory and reputational challenges faced by AI developers in data procurement. With growing discussion around copyright, training transparency, and data provenance in AI systems, pre-digital sources with clear provenance offer a way to mitigate potential compliance risks.


Source: www.golem.de · Published 22 July 2026
Lumi AI News — AI-assisted curation in accordance with Article 50 EU AI Act. Paraphrase and classification by Lumi News Pipeline v1.7.3.

Share on: