Home › Knowledge Hub › Weekly Guidance Watch › AP Handreiking scraping

Regulator guidance · surfaced resource

A regulator’s handbook on scraping by private organisations — and what it says about when it can be lawful

Autoriteit Persoonsgegevens (Netherlands)Handreiking — April 2025 editionDutch only

Almost everything published on scraping in the last two years addresses the training of generative AI models. The question that reaches a data protection officer is often a smaller one — a recruitment team building a candidate list, a sales team enriching a CRM, a research team pulling posts from a public forum. The Dutch supervisory authority has published a handbook aimed at private organisations and individuals doing exactly that, and it may be useful to anyone facing the question.

Published by
The Autoriteit Persoonsgegevens (AP), the Netherlands’ data protection supervisory authority. We have no relationship with the AP and inclusion here is not an endorsement in either direction.
Type
Handreiking — practical guidance (PDF, 427 kB). The AP files it under Persoonsgegevens op internet and Big data en profilering.
Versions & language
April 2025 edition; the AP’s document page is dated 2 April 2025 and carries this edition only. It supersedes the May 2024 first edition. Dutch only — the AP’s own English link for this document returns “this page is not translated”. Note that most secondary commentary still refers to the May 2024 version.
Jurisdiction
Netherlands, under the EU GDPR. It carries no authority in the United Kingdom; see the UK position below.
Primary audience
Private organisations and individuals who scrape, or who buy and use scraped data — and the DPOs and advisers asked to sign it off.
Topic tags
web scraping · legitimate interest · publicly available data · OSINT · AI training data
Availability
Free, on the AP website.

Why it matters

The AP’s headline position is unusually direct for a regulator: “Scraping door private partijen en particulieren is vrijwel nooit toegestaan” — scraping by private parties and individuals is almost never permitted. The handbook exists to help organisations work out whether their own case is one of the narrow exceptions, and the AP’s framing is that lawful scraping has to be tightly targeted rather than broad.

What makes it worth reading outside the Netherlands is its scope. It is not limited to AI training — it also deals with scraping to build lists, to enrich records and to research individuals, which is where, in our experience, most of the practical questions arise. It has never been translated, and we have found little discussion of it in English.

The UK position alongside

The Information Commissioner’s Office has set out its own position on scraping in the narrower context of generative AI training, in its response to the consultation series on generative AI (published 13 December 2024). The ICO treats UK GDPR Article 6(1)(f) legitimate interests as the sole available lawful basis for training generative AI models on web-scraped personal data on current practices, and it expects controllers to “evidence why other available methods for data collection are not suitable”. On the three-part test, the ICO says developers should evidence the likely benefits rather than assume them, that the necessity of scraping is “not a settled issue”, and that scraping for generative AI training is a “high-risk, invisible processing activity” where insufficient transparency will likely fail the balancing test.

The divergence, stated plainly. The AP guidance is broader in scope than the ICO’s — it is not confined to AI training — and it is framed more restrictively. Neither binds the other, and a controller operating in both jurisdictions has to satisfy both. Where an organisation is relying on a legitimate-interests assessment built only against UK guidance, the Dutch handbook is a useful stress test of it.

Two more resources on the same question

The CNIL’s fiche focus on collection by moissonnage (web scraping), published 19 June 2025 and in French, is the operational counterpart: it separates measures it treats as mandatory from additional safeguards, and it is specific in a way that most guidance is not. Among the safeguards is a measure that is easy to miss — using “des pseudonymes aléatoires propre à chaque contenu … et non à chaque identifiant”, a random pseudonym per item of content rather than per identifier, so that scraped material cannot be cross-linked back into a profile. The fiche also covers the opposite side of the problem: how site publishers can protect their own content from being scraped.

The view from the platform’s side

The Concluding joint statement on data scraping and the protection of privacy (28 October 2024) comes at it from the platform’s side. It was produced by the Global Privacy Assembly’s International Enforcement Cooperation Working Group and signed by sixteen authorities including the ICO, following direct engagement with industry. It states that “Personal information that is publicly accessible is subject to data protection and privacy laws in most jurisdictions”, and sets out what organisations holding publicly accessible personal data are expected to do about scraping — rate limiting, bot detection, monitoring for unusual account behaviour, a named team owning the controls, contractual terms where scraping is permitted, and a preference for providing an API so that access stays controlled.

The live one — and a note on it

The European Data Protection Board has adopted Guidelines 03/2026 on web scraping in the context of generative AI, Version 1.0, for public consultation — adopted 7 July 2026, open for comment until 30 October 2026. These are a draft and they may change. They treat Article 6(1)(f) as the basis most commonly available to private entities scraping for generative AI training, and set out mitigating measures that overlap closely with the CNIL’s list — excluding sensitive categories, restricting collection from particular sites, pseudonymisation, rights mechanisms and transparency safeguards. On our reading the two point in the same direction, which means anyone building a scraping assessment now has a published, concrete statement of what the measures look like in practice rather than a reason to wait.

There has already been a great deal written about the draft, much of it by law firms writing on LinkedIn, and we have not tried to add to it. It is listed here simply so that it is easy to find alongside the rest of the reference material.

A Weekly Guidance Watch resource entry, curated by VulaPri. We summarise and link to the original; we do not reproduce or host it. Facts verified against the primary sources on 23 September 2026. Suggest a correction.