Skip to content

AI Scraping of Personal Data: What the GDPR Requires

In a nutshell: Web scraping for AI training regularly captures personal data and remains fully subject to GDPR requirements despite the data’s public availability.

When training data for generative AI systems is scraped from the internet, it regularly includes personal data as well. The GDPR sets clear limits on this practice, which providers and operators of AI systems must comply with.

Web scraping refers to the automated extraction of content from publicly accessible websites, forums, social networks and other online sources. Generative AI models require large volumes of data for their training, which are frequently obtained through this method. Since the internet contains not only purely factual information but also names, photos, contact details, expressions of opinion and other personal information that is freely accessible, this data is regularly captured as part of the scraping process as well. The fact that data is publicly viewable does not change the fact that its processing is subject to the rules of the GDPR.

For controllers who deploy AI systems or commission their development, this gives rise to a duty of review across several data protection principles. First, a legal basis under Art. 6 GDPR is required for processing the scraped personal data, with legitimate interest regularly being relied upon as the basis, although this presupposes a balancing of interests against the rights of the data subjects. In addition, the principles of purpose limitation, data minimisation and transparency come into play, which are particularly difficult to implement when collecting large, unstructured volumes of data from the web, since data subjects generally have no way of knowing that their data is being used for AI training.

For compliance officers at companies that use generative AI applications or train their own models, this means that the origin and lawfulness of training data must be traceably documented and reviewed by providers. This also includes clarifying how data subject rights such as access, erasure or objection can be technically and organisationally implemented with respect to scraped data that has been incorporated into the model, since removing individual data records from an already trained model after the fact is practically almost impossible. When procuring or further developing AI systems, companies should therefore demand contractual assurances from providers regarding data protection-compliant data acquisition and document these as part of a data protection impact assessment.


Source: www.computerweekly.com · Published 19 August 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.8.3.

Share on: