European Data Protection Board official logo, identifying the source of the draft web-scraping guidance
GDPR implementationGDPR and AI web scraping: public data is not blanket permission
European Data Protection Board logo, identifying the source institution. Editorial reference only; no endorsement or affiliation.Image: European Data Protection BoardEDPB reuse conditionsImage fitted for layout; headline added by Securelay. Editorial context, not endorsement.
Primary source EDPB, 8 July 2026: announcement of web-scraping guidance and consultationPrimary source EDPB: Guidelines 03/2026 public consultation

What changed in July 2026

The EDPB announced Guidelines 03/2026 on web scraping for generative AI on 8 July. As checked on 7 September, the guidance is open for consultation, with feedback due on 30 October 2026. It is not final guidance or a replacement for the GDPR. The distinction matters when updating a privacy assessment or presenting the work to a customer.

Public access is not the whole assessment

The EDPB explains that GDPR can apply when scraping involves personal-data processing. Its announcement addresses lawful basis, transparency, purpose limitation and minimisation. It also notes that special-category data needs both an Article 6 basis and an applicable Article 9 exception. The fact that a page is publicly visible is not, by itself, that assessment.

Start with a collection decision you can explain

Our suggested intake record names the intended AI use, source categories, fields to collect, exclusions, retention period and accountable owner. Ask whether the task can use licensed, synthetic or less detailed material instead. Separate access to a website from the rights to reuse its contents: copyright, contractual restrictions and data-protection duties are different questions.

Keep a source record without building another sensitive archive

The EDPB recommends attention to reliable sources, collection timestamps and validation. In an engineering review, decide how to track provenance without retaining unnecessary raw copies in logs. Use a restricted source register and a deletion path. Test how a source exclusion reaches scheduled collection jobs and already-staged data.

Ask the difficult deletion question before training

As a practical design exercise, trace a synthetic record from collection through cleaning, intermediate files and the training dataset. Identify which copies your team can remove and how it would verify that action. Do not equate deleting a source row with removing its influence from a trained model. Require a separate assessment of that problem before promising an outcome.

Where Securelay fits

Securelay's configured detection, redaction and tokenisation can help reduce exposure on integrated data routes. These controls do not establish a lawful basis, grant scraping rights or make a training dataset automatically anonymous. A useful first discussion focuses on the collection boundary, excluded data classes and the evidence required for each approved use.

Put the control on the data path.Discuss an architecture review