US, EU, UK Web Scraping Rules for 2026: EDPB Guidance and Key Cases

Web scraping is generally lawful when you collect public, non-personal data without breaching a contract or hammering someone's servers, but that answer changes fast once personal data, login walls, or copyrighted content enter the picture. The real risk sits in four places: what you access, what you take, what you agreed to, and how hard you hit the target. If your project touches any of those at scale, a legal review before launch is cheaper than one after a cease-and-desist letter arrives.
TL;DR:
- Scraping publicly available data is generally lawful, but risks increase significantly if personal data, login bypass, or copyrighted content are involved at scale.
- In the US, legal exposure mainly stems from contract violations, property claims, and data protection laws, with recent cases narrowing the scope of CFAA liability for public pages.
- The EU and UK treat collection of personal data as always GDPR-regulated, requiring a lawful basis and adherence to data minimization, opt-outs, and retention limits, even on public websites.
- Technical choices such as rate limiting, respecting robots.txt, and using genuine mobile proxies can mitigate legal risks by aligning traffic patterns with normal user behavior.
- Scraping violations often result in civil claims like contract breaches and copyright or database rights infringement, rather than criminal hacking charges, emphasizing the need for thorough legal and technical review beforehand.
Table of Contents
- Is Web Scraping Legal? The Framework That Actually Decides It
- United States: CFAA, Contract Risk, and What Courts Actually Punish
- EU and UK: GDPR, the 2026 EDPB Guidance, and Database Rights
- The Legal Claims Scrapers Actually Get Hit With
- Technical Safeguards That Reduce Legal Exposure
- What the Major Cases Actually Teach You
- A Pre-Launch Legal-Risk Checklist for Your Scraper
- What Fifteen Years Of Watching This Space Has Taught Me
- Where Mobile Proxies Fit Into a Compliant Scraping Setup
- Primary Sources Worth Reading Directly
- Sources
- FAQ
Is Web Scraping Legal? The Framework That Actually Decides It
Most people ask "is web scraping legal" expecting a yes or no. The honest answer is that legality turns on four separate axes, and a project can pass on three of them and still get you sued on the fourth. Understanding these axes matters more than memorizing any single court ruling, because the rulings themselves are just applications of these same variables to specific facts.
Access is the first axis. Did you reach the data by visiting a public page, or did you log in, bypass a paywall, or circumvent a technical block? Courts and regulators treat these very differently. Scraping a public product listing page is a different act, legally, than scraping behind an authentication wall using credentials you were not given.
Content is the second axis. Facts like prices, addresses, and stock levels carry little copyright protection in most jurisdictions. Creative writing, photography, and curated databases carry more. In the European Union and UK, a compiled database can also earn its own sui generis protection separate from copyright, even if none of the individual data points are protected on their own.
Contract is the third axis, and it's the one most scrapers underestimate. A website's terms of service can prohibit automated collection even when the underlying content has no copyright or database-right protection at all. Whether that prohibition is enforceable against you depends heavily on how the terms were presented, which is where the clickwrap versus browsewrap distinction comes in later in this guide.
Conduct is the fourth axis: how your scraper behaves technically. A slow, respectful crawler that mimics normal browsing traffic looks very different in a courtroom than a script firing thousands of requests per second from a single IP, ignoring rate limits, and degrading site performance for other visitors.
One point trips up almost everyone new to this space: "public" does not mean "free of data-protection rules." A LinkedIn profile, a public Instagram post, or a government registry listing a person's name and address is publicly visible, but in the EU and UK it is still personal data under the EDPB's 2026 guidance, and collecting it triggers GDPR obligations regardless of visibility. The US treats this differently in most states, but that gap is narrowing as more states pass their own privacy statutes modeled loosely on GDPR.
A handful of quick markers tend to predict whether a scraping project is low risk or high risk before you write a single line of code:
- Login required to view the data escalates risk immediately, since it usually means you're agreeing to terms and possibly bypassing access controls.
- The data includes names, emails, biometric details, or other identifiers tied to real people shifts the project into data-protection territory, even from public pages.
- The target site has explicit anti-scraping language in its terms raises contract exposure, particularly if you created an account to access the data.
- Your crawler ignores robots.txt or rate limits and causes measurable slowdowns opens the door to server-harm claims.
- You plan to redistribute or resell the scraped dataset raises copyright and database-right exposure on top of everything else.
- The target audience or servers sit in a different country than you do adds a jurisdictional layer, since EU and UK rules can apply to you even if your company is based elsewhere, depending on who the data subjects are.
None of these markers make a project automatically illegal. They tell you where to spend your legal-review time before you scale up.
United States: CFAA, Contract Risk, and What Courts Actually Punish
In the US, the central legal question has been whether scraping violates the Computer Fraud and Abuse Act, a 1986 anti-hacking statute never written with modern data collection in mind. The two cases every US-facing project should understand are hiQ Labs v. LinkedIn and Van Buren v. United States.
In hiQ v. LinkedIn, the Ninth Circuit held that scraping publicly accessible LinkedIn profile pages likely does not violate the CFAA's prohibition on accessing a computer "without authorization," because the data sat behind no login wall and was visible to any visitor. That ruling was a real win for scrapers who stick to public pages. But it wasn't the end of the story: LinkedIn's contract claims against hiQ survived, and the case ultimately settled on those contract grounds rather than ending in a clean scraping victory. The lesson is that beating a CFAA claim doesn't mean beating the lawsuit.
Van Buren narrowed things further in 2021. The Supreme Court held that "exceeds authorized access" under the CFAA applies to accessing files or areas a person had no right to access at all, not to misusing access they were otherwise entitled to. This narrowed reading makes it harder for a website owner to argue that scraping public pages "exceeds authorization" just because the site's terms disapprove of automated collection. Combined, these two rulings mean CFAA claims against public-page scraping are weaker than they were a decade ago.
That doesn't mean scraping in the US carries no risk. It means the risk has shifted from criminal-flavored hacking statutes to civil contract and property claims:
- Breach of contract claims target scrapers who accepted a clickwrap agreement, most often by creating an account, and then violated its no-automation clause.
- Trespass to chattels claims argue that a scraper's traffic imposed a real cost on the target's servers, similar to older cases against unwanted email senders.
- Court injunctions remain the most common remedy sought, not criminal charges. Companies generally want the scraping stopped, not the scraper prosecuted.
- Copyright claims appear when the scraped content includes protected creative work rather than raw facts.
The practical takeaways for anyone building a US-facing scraper are straightforward. Avoid scraping behind a login unless you have explicit authorization from the account holder or the platform, since authenticated scraping reintroduces the exact CFAA exposure that hiQ and Van Buren helped narrow for public pages. Keep your request rate low enough that you can credibly argue you caused no server harm if challenged. And read the terms of service before you scrape, not after a demand letter arrives, because a browsewrap notice buried in a footer link is far weaker evidence of an enforceable agreement than a checkbox you clicked during signup. Quinn Emanuel's analysis of the scraping landscape makes the same point from the defense side: scraping is not per se illegal in the US, but liability is fact-specific, and the facts that matter most are contract terms, access method, and server impact.
EU and UK: GDPR, the 2026 EDPB Guidance, and Database Rights
The European Union and UK add a layer of complexity the US mostly lacks: data-protection law that applies to personal data regardless of whether that data was ever gated behind a login. If a scraped record can identify a living person, GDPR applies the moment you collect it, not just when you publish or sell it.
The EDPB's 2026 guidelines on web scraping confirm this directly for the generative AI context, clarifying that scraping personal data to build or train AI models triggers GDPR obligations just like any other form of personal-data processing. Controllers need a lawful basis to process that data, and for most scraping projects the realistic option is "legitimate interest," which requires passing a documented balancing test weighing the controller's interest against the data subject's rights and reasonable expectations.
The EDPB guidance goes further than past statements by laying out concrete expectations for AI training datasets specifically: limiting collection of sensitive categories of data, documenting the legality assessment before large-scale collection begins, and applying technical mitigations to reduce the risk that a trained model memorizes and later regurgitates identifiable personal information. France's data-protection authority, the CNIL, echoes this in its own focus sheet, recommending that controllers relying on legitimate interest implement push-back mechanisms so individuals can object, maintain exclusion lists for sources or people who opt out, and anonymize or delete collected personal data as soon as it's no longer needed.
Beyond GDPR, European law adds a second protection US law doesn't have in the same form: the sui generis database right under the EU's Database Directive, which protects compilations that require substantial investment to create, gather, or verify, separate from any copyright in the individual entries. Scraping a large chunk of a protected database can trigger liability even when none of the scraped facts are copyrightable on their own. The counterweight is the text-and-data-mining exception under the EU's Copyright in the Digital Single Market Directive, which permits TDM for scientific research and, under certain conditions, for general TDM purposes, unless the rights holder has expressly reserved those rights, often through a machine-readable opt-out.
Regulatory signal to watch: the EDPB's 2026 guidance explicitly ties AI-training scraping to documented legality assessments and technical mitigation, not just a general privacy notice buried on a website. Treat "we posted a privacy policy" as the floor, not the ceiling, of compliance.
A short list of practical safeguards emerges directly from this EU and UK guidance:
- Run and document a legitimate-interest balancing test before large-scale collection of personal data, as the CNIL recommends.
- Honor machine-readable TDM opt-outs rather than treating the absence of a paywall as blanket permission to mine content.
- Limit sensitive-category data in AI training scrapes, since biometric, health, and similar categories carry extra restrictions under GDPR.
- Build a push-back or opt-out channel so individuals can request exclusion from a dataset before or after collection.
- Anonymize or delete personal data on a defined schedule rather than retaining raw scraped records indefinitely.
The Legal Claims Scrapers Actually Get Hit With
"Illegal" is rarely the right word for what happens when a scraping project goes wrong. What actually shows up in a courtroom is one or more specific civil claims, each with its own elements and its own defenses. Knowing which claim applies to your situation tells you exactly what to fix.
- Breach of contract. This claim succeeds when a target site can show you agreed to terms prohibiting automated access, most commonly through a clickwrap agreement where you actively clicked "I agree" during account creation. Browsewrap terms, buried in a footer link nobody clicks, are far weaker evidence of a binding agreement and courts have repeatedly declined to enforce them against parties who never affirmatively accepted them.
- Copyright infringement. This claim targets the content you scraped, not the act of scraping itself. Raw facts like prices and addresses generally fall outside copyright protection, following the reasoning in Feist Publications v. Rural Telephone Service, but product descriptions, articles, photos, and other creative content can be protected. TDM exceptions in the EU can carve out research and certain AI-training uses, but those exceptions come with conditions and opt-out mechanics that vary by member state.
- Database-right infringement. In the EU and UK, extracting or reusing a "substantial part" of a protected database can trigger liability even without any individual copyright violation, if the database owner invested substantially in compiling it.
- Trespass to chattels. This US-rooted claim argues that your scraping traffic imposed a measurable cost or interference on the target's computer systems, similar to the logic used against early spam senders. It requires showing real harm, not mere annoyance, which is one reason polite, rate-limited crawlers rarely face this claim successfully.
- Unjust enrichment. Occasionally pled alongside the other claims, this argues that a scraper profited unfairly from another party's investment in creating or maintaining the scraped content, without paying for that value.
Understanding which of these five claims fits your specific data source and use case is more useful than asking a blanket "is this legal" question, because each one has different defenses and different evidence requirements.
Technical Safeguards That Reduce Legal Exposure
Legal risk in scraping is rarely purely legal. It's often a downstream consequence of technical choices made months before a lawyer ever sees the project. Engineering decisions and legal exposure are more connected here than in almost any other software compliance area.
Rate limiting is the single highest-leverage technical control available. A crawler that spaces out requests, respects concurrency limits, and mimics normal browsing patterns is far less likely to trigger a trespass-to-chattels claim or a server-harm argument, because the core element of those claims is measurable interference with normal operations. Masklabs covers the mechanics of this in more depth in a guide to rate limiting for large-scale automation, but the short version is: slower and steadier beats fast and aggressive, both technically and legally.
Robots.txt and CAPTCHAs deserve a specific caution. Following robots.txt directives and avoiding CAPTCHA circumvention is good practice and generates useful evidence of good faith, but neither one is a legal permission slip on its own, nor does ignoring them automatically create liability. OWASP's material on automated threats frames these as technical controls that inform, but don't determine, the legal analysis. Document any deliberate exception to a robots.txt rule and the reasoning behind it, because that documentation becomes the evidence trail if the decision is ever questioned.
For personal data specifically, minimization and deletion timelines matter more than most engineering teams assume. Collect only the fields you actually need, strip or hash identifiers you don't need for your use case, and set a defined retention window rather than storing raw scraped records forever. This isn't just a GDPR nicety. As datasets age, the argument that you're still using them for a legitimate, documented purpose gets weaker.
For AI training pipelines built on scraped data, a few operational controls carry outsized weight:
- Sample-check your dataset for personal or sensitive data before training rather than after, since removing memorized information from a trained model is far harder than filtering it out upfront.
- Maintain an exclusion list of sources or individuals who have opted out, and check new scrapes against it.
- Pseudonymize identifiers where the training task doesn't actually require a real name, email, or handle.
- Log your legality assessment for the dataset, including which lawful basis you're relying on, so you can show your reasoning if a regulator asks.
Masklabs's own breakdown of building AI training datasets with mobile proxies covers how infrastructure choices intersect with these same operational controls.
Pro Tip: Keep a simple changelog of every technical decision that touches a scraper's behavior, rate limits, excluded domains, robots.txt overrides, and the date each change went live. If a dispute ever surfaces, that changelog often does more to demonstrate good faith than any policy document.
What the Major Cases Actually Teach You
Court rulings and enforcement actions are more useful as pattern libraries than as strict rules, since most scraping disputes settle before trial and never produce a binding precedent. Four situations stand out for the practical lessons they leave behind.
hiQ Labs v. LinkedIn remains the reference case for public-page scraping. The Ninth Circuit's finding that scraping publicly visible profiles likely doesn't violate the CFAA gave scrapers real breathing room on the criminal-hacking-statute front. But the eventual settlement on contract grounds is the part people forget to mention: winning the CFAA argument didn't make hiQ's business viable, because LinkedIn's terms-of-service claim kept the pressure on regardless.
Van Buren v. United States taught a narrower, more technical lesson: "exceeds authorized access" means accessing something you had no right to access at all, not misusing access you were otherwise permitted. That distinction matters for anyone scraping data they can technically reach but that a site's terms say they shouldn't.
BrandTotal's dispute with Meta highlights a different wrinkle: whose computer counts as "accessed" when a browser extension collects data on behalf of a user who's logged into their own account. Cases like this show that authorization questions don't always have a clean answer when third-party tools sit between the user and the platform.
Clearview AI's fines across multiple European data-protection authorities for scraping facial images to build a biometric-matching database illustrate the ceiling of enforcement risk when scraped data is both personal and sensitive. Biometric data sits in a protected category under GDPR, and regulatory commentary has consistently flagged this kind of scraping as high risk regardless of the data's public visibility.
The through-line across all four: authorization, contract terms, and the sensitivity of the data collected matter more than whether the scraping itself was technically difficult to detect.
A Pre-Launch Legal-Risk Checklist for Your Scraper
Run through these questions before your scraper touches production traffic, in roughly this order:
- Is the data personal? If yes, identify your lawful basis (likely legitimate interest) and document a balancing test before collecting.
- Does accessing it require login credentials? If yes, confirm you have explicit authorization; authenticated scraping carries the highest contract and CFAA-adjacent risk.
- Do the site's terms of service prohibit automated access? If yes and you accepted them via clickwrap, treat this as a real contract risk, not a formality.
- Will your request volume measurably affect server performance? If uncertain, start with conservative rate limits and monitor response times.
- Does your target audience or the data subjects sit in the EU or UK? If yes, GDPR obligations likely apply regardless of your own location.
- Are you scraping creative content, or a large slice of a curated database? If yes, review copyright and database-right exposure before scaling up.
- Do you have a documentation trail covering your legality assessment, rate limits, and any robots.txt exceptions? If no, build one before, not after, a dispute arises.
| Checklist trigger | Suggested mitigation |
|---|---|
| Personal data involved | Document legitimate-interest balancing test |
| Login required | Confirm explicit authorization or stop |
| ToS prohibits automation | Escalate to legal review before proceeding |
| High request volume | Apply conservative, documented rate limits |
| EU or UK data subjects | Apply GDPR-consistent safeguards regardless of your location |
| Creative or curated database content | Review copyright and database-right exposure |
If more than two of these trigger simultaneously, especially personal data plus authenticated access, stop and get counsel involved before writing more code. That combination is exactly the fact pattern that turns a side project into a lawsuit.
What Fifteen Years Of Watching This Space Has Taught Me
The gap between what people fear about scraping law and what actually gets litigated is enormous. Most first-time scrapers worry about criminal hacking charges. In practice, the far more common outcome is a cease-and-desist letter over a contract clause nobody read carefully, or a GDPR complaint over data nobody realized counted as "personal."
What gets underestimated is how much of this risk is actually an infrastructure decision in disguise. A scraper that fires from a single datacenter IP at maximum speed looks, to any detection system and any court later reviewing logs, exactly like an attack. A scraper that spreads requests across real, geographically distributed IPs at a human pace looks like normal traffic, because in a meaningful sense, it is. Rotating sessions and rate limiting aren't just anti-blocking tactics. They're evidence, if this ever goes to a dispute, that you weren't trying to overwhelm anyone's systems.
The teams that get this right treat legal review and engineering as one conversation, not two separate departments handing off a finished product. Legal should know your rate limits. Engineering should know which data categories trigger GDPR. That collaboration is worth more than any single clause in a terms-of-service document, because the law here keeps moving, and infrastructure choices made with legal exposure in mind age far better than ones made purely for throughput.
— Jon
Where Mobile Proxies Fit Into a Compliant Scraping Setup
Masklabs is not a law firm, and nothing here replaces a real legal review, but the infrastructure you choose does shape how defensible your scraping operation looks. The provider offers real carrier IP addresses across multiple US cities, with rotating or sticky sessions depending on whether your project needs a consistent identity for a session or fresh IPs on every request.

That distinction matters for the rate-limiting and server-harm discussion above: distributing requests across genuine mobile-carrier IPs, rather than a single datacenter block, produces traffic patterns that look like ordinary mobile browsing rather than a coordinated scrape. Geolocation targeting also helps with the jurisdictional questions covered earlier, since projects that need city-specific US data can route requests through the actual region rather than approximating it. None of this replaces documenting your lawful basis or reviewing a site's terms of service. If you want to see how Masklabs's mobile proxy infrastructure fits into a project you're already scoping, start with a trial run against your actual target sites before committing to a data plan.
Primary Sources Worth Reading Directly
Court opinions and regulator guidance change faster than any blog post can track, so bookmark the primary sources rather than relying on secondhand summaries. The EDPB's 2026 guidelines on web scraping is the current reference for GDPR and AI training obligations in the EU. The CNIL's focus sheet on legitimate interest gives concrete, operational recommendations. The Ninth Circuit's opinion in hiQ v. LinkedIn remains essential reading for US CFAA questions, and Quinn Emanuel's analysis offers a clear practitioner's summary of the broader legal landscape. Read the originals before making a compliance decision that depends on their exact wording.
This article is general information, not a substitute for advice from a qualified lawyer. Consult a qualified legal professional about your own circumstances before acting on anything here.
Sources
- hiQ Labs, Inc. v. LinkedIn Corp. (9th Cir. opinion)
- EDPB Guidelines 03/2026 on web scraping in the context of generative AI
- The legal basis of legitimate interest: focus sheet on the measures to implement in the case of data collection by web scraping | CNIL
- The Legal Landscape of Web Scraping (Quinn Emanuel)
FAQ
Which websites are legal to scrape?
Public pages that don't require login, don't contain personal data, and aren't protected by explicit contract terms carry the lowest legal risk, but no website is universally "safe" to scrape without checking its specific terms and content type.
Can web scraping be detected?
Yes. Sites commonly detect scraping through IP reputation, request patterns, and behavioral signals like unnatural click speed, which is why rate limiting and realistic traffic patterns, such as those from genuine mobile carrier IPs, reduce both detection and legal exposure at the same time.
Is AI scraping illegal?
Scraping to build AI training data is not automatically illegal, but the EDPB's 2026 guidelines treat it as high risk when personal data is involved, requiring a documented lawful basis and mitigation measures like exclusion lists and anonymization.
Can ChatGPT do web scraping?
ChatGPT itself does not scrape live websites on its own; separate scraping tools and infrastructure, including proxy services, collect the raw data before it's used for training or retrieval, and that collection step is where the legal questions in this guide actually apply.