EDPB Opens Public Consultation on Web Scraping Guidelines for Generative AI: What Businesses Need to Know
On 7 July 2026, the European Data Protection Board (EDPB) adopted the draft Guidelines 03/2026 on web scraping for generative AI, open for public consultation from 8 July to 30 October 2026, a dedicated instrument on the practice behind large-scale AI training: harvesting data from the open web. Whether a business scrapes itself, outsources it, or simply buys a ready-made dataset, the draft Guidelines set out in concrete terms what compliance should look like.
Who is covered
The Guidelines apply to private-sector organizations that scrape personal data from external sources to train or fine-tune generative AI models – whether in-house, via a contractor, or by acquiring a dataset already assembled by another organization. They exclude data brokers who merely resell scraped datasets without training models, an organization's processing of its own data, and scraping by public bodies. Those situations remain subject to data protection law generally; they are simply outside this instrument's scope.
Scraper’s status under GDPR
The entity doing the scraping is not automatically the controller. A scraper acting on an AI developer's documented instructions is likely a processor, with the developer as controller. Where a developer reuses a dataset another organization already scraped, each is responsible only for its own processing – the original scraper is not liable for later reuse. Joint controllership arises when both parties jointly decide the purposes and means, for example agreeing together on the collection criteria for a dataset one will build for the other to train on.
This classification decides who signs the data processing agreement, handles data subject requests, and bears accountability risk if the dataset proves unlawfully collected.
Three GDPR Principles in Focus
What the draft Guidelines add is the EDPB's recommended reading of how each core GDPR principle applies to scraping:
- Transparency: individual notice to every data subject is often impossible or disproportionate, and Article 14(5)(b) GDPR may excuse it – but only after balancing the number of data subjects, the age of the data, and existing safeguards. Even then, a public notice is mandatory, built to a specific checklist:
- a detailed, plain-language Privacy Notice, easy to find;
- the fullest practicable disclosure of sources – ideally a searchable list of domains and URLs scraped, with collection dates;
- the categories of data collected and the purposes and legal basis of processing;
- a working mechanism for individuals to exercise their data subject rights.
- Data minimization: before scraping – precise collection criteria, data mapping, filters excluding categories like financial or location data, exclusion of sites structurally holding sensitive data, and respect for robots.txt, ai.txt and CAPTCHA. During and after collection – syntax-based filtering, synthetic data substitution, anonymization or pseudonymization.
- Accuracy: prefer reliable, maintained sources, timestamp collection dates, and validate data before training – extending to the accuracy of personal data the trained model later outputs.
Legitimate Interest Test
Consent is rarely workable for web scraping: there is no direct relationship with the individual, and publishing data online is not consent to its reuse. Legitimate interest under Article 6(1)(f) GDPR is therefore the basis most private organizations will rely on, subject to three cumulative conditions.
- A genuine, lawfully articulated interest – developing or improving an AI model can qualify, including a general-purpose model with no fixed end use yet.
- Necessity – collection criteria must be narrowed to what the purpose requires; indiscriminate crawling weakens the case, as does ignoring less intrusive alternatives like synthetic or pseudonymized data.
- A balancing test – weighing the controller's interest against the data subject's rights, considering data sensitivity, the scale and duration of collection, and how hard it is for individuals to object once a model is trained.
- where and in what context the data was originally published;
- how public the source really is (an open blog differs from a login-gated forum);
- whether the site imposes technical barriers to automated collection, e.g., robots.txt, ai.txt or CAPTCHA;
- whether the individual could reasonably expect their data to train an AI model – rather than simply be read by other people.
- the characteristics of the data subject, such as being a minor or a public figure – Article 6(1)(f) GDPR itself singles out children's interests as carrying extra weight in this balance.
Where the balance tips against the controller, mitigating measures can restore it: excluding high-risk or login-gated content by default, imposing time limits, publishing an updated list of scraped sources, offering a right to object or an opt-out registry, deleting or anonymizing data promptly, and guarding against model memorization and regurgitation. Adopting none of these – as with a voice-cloning tool trained without safeguards – defeats reliance on legitimate interest outright.
Special categories of data nuances
Scraping of special categories of personal data – such as racial or ethnic origin, political opinions, or health data – is prohibited without an Article 9(2) derogation. Since scraping often cannot avoid such data, the EDPB applies the CJEU's reasoning in GC and Others (C-136/17), originally developed for search engines: the Article 9 prohibition binds a controller only within its own responsibilities, powers and capabilities.
This applies only where the processing resembles a search engine's activity, collection of special-category data is incidental and unintended, it is genuinely hard to screen out in advance, and the controller has implemented safeguards across the data lifecycle – pre-collection filtering, prompt deletion once identified, resistance to extraction during model development, and continuous output monitoring after deployment, with model unlearning should also be considered as necessary.
REVERA recommendations
- Map data sources and the legal basis before scraping begins, not retrospectively – the EDPB treats data minimization as a design obligation.
- Build exclusion logic honoring robots.txt, ai.txt and CAPTCHA signals, and exclude by default sites known to hold sensitive or minors' data.
- Fix the controller/processor/joint-controller allocation in writing before engaging a scraping contractor or purchasing a dataset, and mirror it in the data processing agreement.
- Prepare an Article 14 public notice covering sources, data categories, crawler characteristics and data subject rights, and keep it current.
- Document the necessity and balancing-test analysis as you go – Article 5(2) accountability means showing the reasoning, not just the outcome.
- Where special categories of data cannot realistically be excluded, evidence safeguards across the full lifecycle – collection filters, prompt deletion, training-stage testing and output monitoring – not just a single control point.
Taken together, the Guidelines put the entire AI development lifecycle under GDPR control – from choosing data sources, through collection and storage, to the measures that stop a trained model from reproducing memorized personal data. Compliance at only one point in that chain will not satisfy the EDPB.
Because controller status turns on who actually sets the purposes and means, agreements with scraping contractors and dataset suppliers should fix data sources, collection and exclusion criteria, audit rights, and warranties on how a purchased dataset was assembled. Since the original scraper is not liable for a downstream developer's reuse of the data, the contractual paper trail becomes the developer's primary evidence of accountability if a regulator asks how a training dataset was built.
Timeline
These Guidelines are still a draft. The EDPB adopted them on 7 July 2026 for public consultation, not in final form: they are not legislation in their own right but the Board's non-binding recommendations on how the existing GDPR should be read and applied to web scraping, and the text may still change before it is finalized.
The consultation runs from 8 July to 30 October 2026, with comments submitted via the EDPB's online form and, absent an opt-out, published with the submitter's name, sector and country. Businesses grappling with the tension between large-scale training data needs and the Guidelines' minimization and mitigating-measures expectations have a limited window to raise implementation concerns before the text is finalized.
The Arbitration & IT Disputes practice advises AI developers, scraping contractors and businesses deploying generative AI tools on GDPR-compliant data sourcing, legitimate-interest and balancing-test assessments, and contractual structuring across the scraper-developer-deployer chain.
Contact our lawyer to learn more
Contact a lawyer