EDPB: draft guidelines on web scraping for generative AI
The European Data Protection Board (EDPB) adopted draft Guidelines 03/2026 on web scraping of publicly available online data for training generative AI models – the first comprehensive GDPR interpretation of this issue. According to the EDPB, consent under Article 6(1)(a) GDPR is generally not a viable legal basis for scraping at scale; the main route remains legitimate interest (Article 6(1)(f) GDPR), which must pass a strict three-part test (a legitimate purpose, necessity, and balancing against data subjects' interests). The guidelines also expect technical and organisational measures – respecting robots.txt and ai.txt files, filtering out special categories of data, publishing a detailed privacy notice, and providing mechanisms for data subjects to exercise their rights before collection. Comments may be submitted until 30 October 2026.