What changed in July 2026
The EDPB announced Guidelines 03/2026 on web scraping for generative AI on 8 July. As checked on 7 September, the guidance is open for consultation, with feedback due on 30 October 2026. It is not final guidance or a replacement for the GDPR. The distinction matters when updating a privacy assessment or presenting the work to a customer.
Public access is not the whole assessment
The EDPB explains that GDPR can apply when scraping involves personal-data processing. Its announcement addresses lawful basis, transparency, purpose limitation and minimisation. It also notes that special-category data needs both an Article 6 basis and an applicable Article 9 exception. The fact that a page is publicly visible is not, by itself, that assessment.
Start with a collection decision you can explain
Our suggested intake record names the intended AI use, source categories, fields to collect, exclusions, retention period and accountable owner. Ask whether the task can use licensed, synthetic or less detailed material instead. Separate access to a website from the rights to reuse its contents: copyright, contractual restrictions and data-protection duties are different questions.
Keep a source record without building another sensitive archive
The EDPB recommends attention to reliable sources, collection timestamps and validation. In an engineering review, decide how to track provenance without retaining unnecessary raw copies in logs. Use a restricted source register and a deletion path. Test how a source exclusion reaches scheduled collection jobs and already-staged data.
Ask the difficult deletion question before training
As a practical design exercise, trace a synthetic record from collection through cleaning, intermediate files and the training dataset. Identify which copies your team can remove and how it would verify that action. Do not equate deleting a source row with removing its influence from a trained model. Require a separate assessment of that problem before promising an outcome.
Where Securelay fits
Securelay's configured detection, redaction and tokenisation can help reduce exposure on integrated data routes. These controls do not establish a lawful basis, grant scraping rights or make a training dataset automatically anonymous. A useful first discussion focuses on the collection boundary, excluded data classes and the evidence required for each approved use.
