Understanding the Issue
Organizations developing generative AI models have been scraping personal data from public websites without proper legal basis or transparency. The European Data Protection Board (EDPB) issued guidelines in July 2025, highlighting a systemic compliance failure: treating publicly accessible data as fair game for AI training, ignoring GDPR requirements.
Companies use automated tools to extract data from websites, social media, and forums, aggregating it into training datasets. This often includes special categories of personal data, processed without informing data subjects. They assume legitimate interest, but this doesn't hold up under scrutiny. The scale of these operations makes individual notification "impossible or requiring excessive effort," which they use as an excuse to skip transparency obligations.
This isn't an isolated incident. It's an industry-wide practice that the EDPB is now addressing.
Regulatory Timeline
The timeline shows how this issue escalated:
- Pre-2025: Organizations routinely scrape web data for AI training, assuming publicly accessible information is exempt from GDPR.
- 4 September 2025: Court of Justice of the EU issues ruling in C-413/23 P EDPS v SRB, clarifying when data qualifies as anonymous.
- 8 July 2025: EDPB adopts guidelines on web scraping for generative AI, establishing clear compliance requirements.
- Consultation period through 30 October 2026: Stakeholders can comment before guidelines are finalized.
The gap between technology deployment and regulatory clarity allowed organizations to make assumptions that are now being corrected.
Compliance Failures
Transparency Failures
Controllers assumed that because individual notification was difficult at scale, they could skip it entirely. They failed to implement alternative transparency mechanisms and didn't maintain accessible documentation of their processing activities. Some didn't even acknowledge that web scraping constituted personal data processing subject to GDPR.
Purpose Limitation Violations
Organizations collected data for one purpose and repurposed it for AI model training without a valid legal basis. They didn't assess whether the original collection context made AI training a compatible further processing activity.
Data Minimisation Gaps
Scraping operations pulled entire datasets without filtering for relevance. Controllers didn't implement technical measures to exclude unnecessary data, contradicting Data Minimisation requirements.
Special Categories Processing Without Valid Exceptions
Web scraping captures special categories of personal data, like health information or political opinions. Controllers processed this data without establishing both a lawful basis under Article 6 and an exception under Article 9(2) of the GDPR.
Accuracy and Data Quality Controls
Organizations didn't validate scraped data before using it in training. They didn't timestamp collection to track data currency or implement mechanisms to detect and correct inaccuracies.
GDPR Requirements
The GDPR establishes specific requirements for web scraping for AI:
- Article 6 (Lawfulness of Processing): Requires a valid legal basis. The EDPB guidelines clarify that legitimate interest can apply, but only after a proper balancing test.
- Article 5(1)(b) (Purpose Limitation): Data collected for one purpose must not be further processed incompatibly. You need to assess compatibility or establish a new legal basis.
- Article 5(1)(c) (Data Minimisation): Collect only data that's adequate, relevant, and limited to what's necessary.
- Article 9 (Special Categories): Prohibits processing special categories of personal data without a lawful basis and a specific exception.
- Article 5(1)(d) (Accuracy): Requires reasonable steps to ensure data accuracy.
- Articles 13-14 (Transparency): Require informing data subjects about processing, even if direct notification is impossible.
Action Items for Your Team
Conduct a Web Scraping Inventory
Document every instance where your organization collects data through automated means. Identify what data you're collecting, the original purpose, your processing purpose, the legal basis, and whether special categories are involved.
Implement Pre-Collection Filtering
Configure your collection tools to exclude unnecessary data types. If training a language model for technical documentation, exclude irrelevant data like forum posts on health conditions.
Establish Source Reliability Criteria
Define what "reliable" means for your use case. Consider factors like clear terms of service, data quality controls, and whether the source is designed for public access.
Record Collection Timestamps
Implement automated timestamping for all scraped data to support accuracy requirements and refresh datasets when source data changes.
Build Special Categories Detection
Implement technical measures to identify and handle special categories of personal data. Options include keyword detection, pattern matching, or manual review.
Create Alternative Transparency Mechanisms
Establish public-facing documentation about your web scraping practices. Explain what data you collect, how you use it, and how individuals can exercise their rights.
Reassess Your Legitimate Interest Claims
Document your balancing test for legitimate interest. Consider the data subject's reasonable expectations, the nature and sensitivity of the data, and whether less intrusive means can achieve your purpose.
Implement Data Subject Rights Procedures
Develop processes to handle Data Subject Access Requests, erasure requests, and objections, even when data is embedded in trained models.
The EDPB guidelines are open for consultation through 30 October 2026. Your feedback can shape the final requirements, but the core principles are clear now. Treat web scraping for AI training as personal data processing subject to GDPR.





