Europe's privacy regulators have told AI companies how web scraping must fit inside the GDPR. At its July plenary, announced by the European Data Protection Board on July 8, 2026, the EDPB adopted Guidelines 03/2026 on web scraping in the context of generative AI and opened a public consultation running until October 30, 2026.
What exactly was adopted?
Two documents, plus a finalization. According to the EDPB's official news release from Brussels dated July 8, 2026, the Board adopted guidelines on anonymisation and guidelines on web scraping in the context of generative AI during its latest plenary, and also adopted the final version of its guidelines on processing personal data through blockchain technologies. The web scraping guidelines were simultaneously published for consultation, with the feedback window running from July 8 to October 30, 2026, per the EDPB's consultation page.
The pairing is deliberate. Generative AI models are trained on data scraped from the web at scale, and much of that data contains personal information; the anonymisation guidelines determine when scraped data stops being personal data, while the scraping guidelines govern the collection itself. The EDPB notes its new anonymisation guidance takes into account the ruling of the Court of Justice of the EU in case C-413/23 P EDPS v SRB of September 4, 2025, tying the framework to binding case law rather than fresh invention.
Why does this matter for AI companies?
Because scraping is how training data gets made. Every large language model with web-scale corpora depends on harvesting pages that include profiles, posts, photographs, and behavioral traces of identifiable people. Under the GDPR, collecting personal data requires a lawful basis, purpose limitations, and respect for data-subject rights — obligations that sit awkwardly with a practice that aggregates billions of pages in bulk. Formal EDPB guidelines give national data protection authorities a common reference for enforcement, which historically means investigations follow the guidelines' contours.
The timing also lands mid-cycle for the AI Act. Transparency obligations under that regulation take effect from August 2, 2026, so providers now face privacy rules on training-data collection stacking on top of disclosure duties on model outputs — two regimes, one pipeline. Companies that treated data sourcing as settled law will need to re-examine it, and companies that already document lawful basis per dataset will find the guidelines a map of what regulators expect that documentation to contain.
What happens between now and October 30?
Consultation, then revision, then final adoption — the EDPB's standard cadence. Stakeholders can submit comments through the form on the consultation page until October 30, 2026, and the Board states that submitted comments are published on its website after screening. Final versions typically follow a subsequent plenary, as happened with the blockchain guidelines finalized at this same session.
For AI developers, the practical reading is straightforward: the era of arguing that scraping is a legal gray area in the EU is closing. The guidelines are not yet final, but their direction — scraping personal data is a regulated processing activity, and anonymity claims must survive the EDPB's own tests — is now written down by the authority that coordinates every national privacy regulator in the bloc.

