Introduction
The surge in the popularity of Artificial intelligence has created a new privacy concern. All the Large Language models (LLMs) need a starting point, and irrespective of the industry that starting point is always the open web. On the internet the LLMs often crawl through forums, news archives, product reviews, personal blogs, social media posts to scrape and learn through all the information present in these. The justification these companies have used historically is simple i.e. ‘the data was public, so we are free to use it’.
Regulators across Europe spent the past two-years discussing whether this justification holds up under the GDPR or not. On 7th July 2026 the European Data Protection Board (EDPB) adopted its Guidelines 03/2026 on web scraping in the context of generative AI. This law ends up closing the “it was public” defence used by AI companies for good.
Why publicly available doesn’t mean freely usable?
The GDPR wasn’t drafted to address the questions around the concept of “public” vs “private” data. It is only concerned with the fact whether the information is related to an identified or identifiable natural person. A LinkedIn profile, a product review signed with a real name, a forum post containing an email address: all of this is personal data the moment someone can be identified from it, regardless of the fact how easy it was to find it.
The new EDPB guidelines makes this explicit. EDPB confirms that GDPR will be governing all the activities relating to web scraping if it involves extracting, cleaning, structuring, or storing personal data. In case of datasets that are mixed with personal and non-personal information, the regulation will be applicable to the extent whatever personal data sits inside it.
This is not a new approach to deal with generative AI. The same logic was used against companies scraping data for entirely different purposes. The Dutch DPA imposed a fine of €30.5 million on against facial recognition firm Clearview AI in 2024, for creating an unauthorized biometric database. Clearview’s facial recognition database was built by harvesting billions of photos from social media without the user’s consent, without providing any proper notice and without providing any means to access their data upon request. Authorities in Italy, France, Germany, Greece, and the UK reached similar conclusions independently.
When does Scraped data becomes Personal data
Not everything being scrapped up in a crawl counts as personal data, but the bar is lower than most people expect. Personal data may include simple things such as a name attached to an opinion, username that can be traced back to a real identity elsewhere, a combination of job title, employer, and city etc. Any of these markers can make a person identifiable “considering all the means reasonably likely to be used” as per GDPR’s own test under Article 4(1) and Recital 26.
The key issue faced by AI developers is the scale. A single scrap of the webpage is easy to assess. However, this same exercise becomes difficult if it is to be done for billions of pages. EDPB acknowledges that organisation will find it difficult to figure out whether and what kind of personal data is present in a dataset and to whom exactly this data belongs to.
But the board doesn’t treat this difficulty as an excuse, it stressed on the fact that the organizations still bear responsibility for demonstrating compliance under the accountability principle in Article 5(2) GDPR. AI systems have the ability to memorise and later reproduce personal data drawn from the training sets as a response. The risk doesn’t end at collection but it persists inside the model and has the potential to resurface the moment the model is prompted in a specific way.
A separate complication also involves where special category of data is involved under Article 9, such as health information, political opinions, sexual orientation, religious beliefs, and similar sensitive categories, which require extra legal grounds beyond an ordinary lawful basis. Scraped web text may routinely contain this kind of information invariably buried inside forum posts, social media threads with no realistic way to filter such information out in advance.
Possible Lawful Bases for AI Training
Article 6 of the GDPR lists six lawful bases for processing personal data these are consent, contract, legal obligation, vital interests, public task, and legitimate interests. If we apply these lawful bases to the context of web scrapping then most of them fall away immediately. For example- it is impossible to take consent from billions of unknown individuals across the open web. Other bases such as contract, legal obligation, vital interests and public tasks do not describe the relation between an AI developer and a random person whose blog got scraped.
So, the only practical lawful bases available to AI companies would be legitimate interest. Both the EU and UK regulators have also come to the same conclusion. The UK’s Information Commissioner’s Office (ICO), after a lot of deliberations concluded that the legitimate interest is practically the only available lawful basis for training generative AI on web-scraped personal data given the current industry practices. The EDPB's parallel Opinion 28/2024 on AI models built the analytical framework that the 2026 scraping guidelines now apply directly to the collection stage of the AI pipeline.
Legitimate Interests and the Three-part Balancing Test
Relying on legitimate interests is not a formality. AI companies ought to perform a balancing test. Both >the EUUK regulators have described this test in a near identical terms:
- Purpose Test: The companies need to find out is there a genuine, specific and lawful interest being pursued (commercial product development, research, fraud detection, and so on), rather than a vague purpose. Organisations need to be specific about the purpose even if they are developing general-purpose AI models.
- Necessity Test: Next, Companies need to assess whether such large-scale scrapping is required or not?
- Balancing Test: Do the rights and interest of the individuals whose data was scraped outweigh the organisation’s interest in using it.
Regulators have stated that the third-step is where most of the AI-web scraping activities struggle. Legal analysis of the ICO’s position notes, notes that web scraping activities are likely to fail the balancing test in majority of the instances because of the fact that there is lack of transparency between the data subjects and the company. What that means is data subjects have no way of knowing that their data was used and therefore they cannot exercise their rights over it.
On the contrary, EDPB’s 2026 guidelines treat this as an operational requirement rather than just a principle to be followed. EU regulators call for the data to be minimised even before the crawl is supposed to happen, not after. Reports on the guidelines state that companies are advised to apply legal-compliance filters ahead of collection, reversing the industry-standard sequence of crawling first and filtering for quality and legality afterward.
The Transparency Problem with Indirect Collection
Article 14 GDPR requires controllers to inform individuals when personal data is collected about them from a source other than the individuals directly, which describes essentially all scraped training data. In a practical sense, telling billions of anonymous internet users that their old post is now a part of a AI training machinery is not a feasible in the sense of individual notice.
This invisible processing is a problem that the regulators have ran into again and again. This is the same concern that was reflected in Open AI’s violation in Italy. The Garante (Italian Data Protection Authority) had fined Open AI €15 Million for processing personal data to train ChatGPT without first identifying an adequate legal basis, and that this obligation should have been adhered to before the chatbot's public release meaning a that a valid basis should have been in place at the very start of processing, not retrofitted afterward. On top of the fine regulators also ordered Open AI to conduct a six-month public communication campaign explaining how ChatGPT collects both users and non-users and what rights people have to object, rectify, or delete it. So, transparency was treated as remedy and not just as a compliance checkbox.
Data Subject Rights in an AI context
Exercising Data subject rights against a trained AI model is technically difficult. One cannot simply delete a single data point from model's parameters the way you would delete a row from a database, because, as the EDPB has noted, training data can remain effectively "absorbed" into the model's mathematical structure, these data sets no longer remain retrievable long after original data set is gone. This does not mean that AI companies can disregard the rights, they are expected to develop build deletion retraining and objection-handling process into the model lifecycle from the start rather than treating it as a separate problem to solve if and when the complaint arrives.
What should AI Developers do before scraping?
Given the accountability principle in Article 5(2), the burden sits squarely with the organization doing the scraping. Before a large-scale crawl begins, that typically means:
- A documented legitimate interest assessment covering the full three-part test
- A data protection impact assessment (DPIA), given the scale and risk profile of the processing
- A defined data minimization strategy applied at collection, not after the fact
- A plan for indirect transparency (privacy notices, public statements, opt-out mechanisms)
- Filtering methods for special category and clearly excessive personal data
- A documented controller/processor analysis, since the EDPB has clarified that the entity doing the scraping is not automatically the GDPR controller, that role depends on the actual facts of who determines the purpose and means of processing¹¹
- A retention and deletion policy addressing what happens when a data subject objects or requests erasure
Conclusion
The era for AI companies using “it was already public” as a legal shield is over. With the EDPB’s 2026 scraping guidelines, its 2024 opinion on AI models, and actions taken against Clearwater AI and Open AI have sent a strong consistent message that publicly accessible does not mean freely processable, and the burden of proving otherwise sits with whoever runs the crawl. EDPB’s guidelines remain open for public consultation until 30 October 2026, which makes now the moment for businesses to get their documentation in order not after the framework is finalized.
We at Data Secure (Data Privacy Automation Solution) DATA SECURE - Data Privacy Automation Solution can help you to understand Privacy and Trust while lawfully processing the personal data and provide Privacy Training and Awareness sessions in order to increase the privacy quotient of the organisation.
We can design and implement RoPA, DPIA and PIA assessments for meeting compliance and mitigating risks as per the requirement of legal and regulatory frameworks on privacy regulations across the globe especially conforming to GDPR, UK DPA 2018, CCPA, India Digital Personal Data Protection Act 2023. For more details, kindly visit DPO India – Your outsourced DPO Partner in 2025 (dpo-india.com).
For any demo/presentation of solutions on Data Privacy and Privacy Management as per EU GDPR, CCPA, CPRA or India DPDP Act 2023 and Secure Email transmission, kindly write to us at info@datasecure.ind.in or dpo@dpo-india.com.
For downloading the various Global Privacy Laws kindly visit the Resources page of DPO India - Your Outsourced DPO Partner in 2025
We serve as a comprehensive resource on the Digital Personal Data Protection Act, 2023 (Digital Personal Data Protection Act 2023 & Draft DPDP Rules 2025), India's landmark legislation on digital personal data protection. It provides access to the full text of the Act, the Draft DPDP Rules 2025, and detailed breakdowns of each chapter, covering topics such as data fiduciary obligations, rights of data principals, and the establishment of the Data Protection Board of India. For more details, kindly visit DPDP Act 2023 – Digital Personal Data Protection Act 2023 & Draft DPDP Rules 2025
We provide in-depth solutions and content on AI Risk Assessment and compliance, privacy regulations, and emerging industry trends. Our goal is to establish a credible platform that keeps businesses and professionals informed while also paving the way for future services in AI and privacy assessments. To Know More, Kindly Visit – Your Trusted Partner in AI Risk Assessment and Privacy Compliance | AI-Nexus