Try 500 MB of US mobile proxy data free for 30 days.Start free trial
September 27, 20269 min read

Developers: 7 Robots.txt Checks to Avoid Legal Risk in AI Training

Isometric robots.txt access decision illustration

Robots.txt is a voluntary site instruction under RFC 9309: respectful scrapers check it before crawling, but it is not an access authorization or a security barrier. Honoring it means fetching the file, parsing its rules correctly, and building fallback logic for cases where the file returns an error or times out. It does not replace rate limiting, and it does not grant you legal permission to take data you would not otherwise be allowed to access.


TL;DR:

  • Robots.txt is voluntary and only a courtesy signal; ignoring it does not grant legal access or provide security.
  • Proper handling requires fetching, parsing, and caching the file with RFC-compliant libraries, and testing against percent-encoding and case variations.
  • Disallow and Allow rules resolve by the longest path match, and wildcards only operate within single path segments, not as full regexes.
  • Server responses with 4xx indicate the file is unavailable and may be ignored, while 5xx responses mean the file is unreachable and should block crawling until resolved.
  • Large-scale scraping for AI training remains legally uncertain, and respecting rate limits, licensing, and personal data restrictions is essential.

Table of Contents

What robots.txt and RFC 9309 actually define

RFC 9309 standardizes the Robots Exclusion Protocol: a plain-text file at the root of a domain that tells crawlers which paths they may or may not fetch. The file lives at /robots.txt, is served as UTF-8 text/plain, and applies to one origin at a time. Inside it, rules are organized into groups, each starting with one or more User-agent lines followed by Allow and Disallow rules. When multiple groups could apply to a crawler, RFC 9309 standardizes matching so the most specific user-agent group wins, not the first one listed.

The protocol exists to manage crawl budget and indexing preferences, letting site owners steer bots away from duplicate pages, admin panels, or infinite parameter combinations. It was never built as an access control mechanism. A path disallowed in robots.txt can still be requested by any client that ignores the file, and as Google's own documentation notes, a blocked page can still be indexed if other signals point to it.

How to fetch, cache, and parse robots.txt reliably in code

Getting robots.txt handling right in production code comes down to a handful of concrete checks, most of which naive scrapers skip.

  1. Fetch the file from /robots.txt at the target origin, following up to five redirects before giving up.
  2. Check the response code: a 4xx means the file is unavailable and you may proceed; a 5xx means it is unreachable and you should hold off.
  3. Cache the parsed rules for no longer than 24 hours unless the file itself is unreachable, in which case shorter retry intervals make sense.
  4. Use an RFC-compliant parser rather than hand-rolled string matching, since edge cases in wildcards and encoding break custom logic quickly.
  5. Confirm the file is plain UTF-8 text: files saved or edited through word processors sometimes introduce byte-order marks or smart quotes that break parsing.

Skipping any of these steps tends to produce scrapers that either crawl paths they shouldn't or, just as often, block themselves from content that was never actually restricted.

Common directives and parsing gotchas

Most robots.txt files use a small set of directives, but the matching logic around them causes more bugs than the syntax itself.

  • User-agent matching determines which group of rules applies to your crawler, and when several groups could match, the most specific user-agent string takes priority over a wildcard * group.
  • Allow and Disallow are resolved by the longest matching path prefix, not by which line appears first, so a short Disallow: / can be overridden by a longer, more specific Allow: /public/.
  • Wildcards (*) and end-of-string anchors ($) let you match patterns, but they only work within a single path segment the way the spec defines them, not as full regular expressions.
  • Percent-encoding and case sensitivity trip up naive implementations: /Café and /caf%C3%A9 may need to be treated as the same path, and comparisons are case-sensitive by default.

Pro Tip: Test your parser against percent-encoded paths and mixed-case URLs before trusting it in production, since substring matching alone will silently misclassify both.

How robots.txt availability and redirects affect access decisions

Server responses when fetching robots.txt carry specific meaning under the spec, and getting this wrong leads scrapers to either overreach or stall unnecessarily. Per RFC 9309, a 4xx response means the file is "unavailable," and a crawler may treat the site as having no restrictions. A 5xx response means the file is "unreachable," and the spec says the crawler must assume complete disallow until it can fetch the file successfully.

Redirects add another layer: RFC 9309 allows following up to five redirect hops when fetching robots.txt, and stopping after that treats the file as unreachable. For transient failures, a scraper that backs off exponentially, temporarily lowers its concurrency, and logs the failure for a human to check tends to stay compliant without grinding to a halt over a single flaky response. Treating every error identically, whether it is a 404 or a 503, is one of the more common mistakes in scraper design.

Robots.txt response and retry decision flow

Legal and ethical considerations for scraping and AI training data

The Robots Exclusion Protocol is voluntary by design: nothing in RFC 9309 creates legal authorization to access content, it only describes a courtesy convention that well-behaved crawlers follow. Whether large-scale scraping for AI training crosses a legal line is still unsettled. The U.S. Copyright Office is actively studying the legal issues around scraping and AI training data, and its ongoing work reflects how unresolved the question remains in the United States.

In practice, that uncertainty should shape how teams operate, not just what they read. For large-scale or commercial data collection, a direct agreement or licensing arrangement with the site owner is a more durable foundation than relying on robots.txt silence as consent. Respecting rate limits and avoiding personal or sensitive data protects both the target site and your own project. Site owners, meanwhile, should avoid listing sensitive paths in robots.txt at all, since the file is public and readable by anyone, including the people it is meant to deter.

Legal and ethical considerations for scraping and AI training data — overview diagram

A compact developer checklist: implement, test, and operate scrapers respectfully

Before a scraper goes into production, run it through a short list of checks that catch most of the mistakes covered above.

  1. Fetch and parse robots.txt with an RFC-compliant library, never a hand-written regex matcher.
  2. Cache parsed rules and refresh them on a defined interval, generally no longer than 24 hours.
  3. Build unit tests that simulate 4xx and 5xx robots.txt responses, plus redirect chains up to five hops.
  4. Include test cases for percent-encoded and mixed-case URLs to confirm the longest-match rule behaves correctly.
  5. Apply conservative rate limits and exponential backoff, especially after a 429 or 403 response.
  6. Watch for CAPTCHA challenges as a signal to slow down rather than a puzzle to defeat programmatically.
  7. Set a clear threshold for when repeated failures should pause the crawl and route to a human for review, rather than retrying indefinitely.

Rate limiting deserves particular attention since it affects both compliance and reliability. Our post on how rate limiting affects large-scale web automation walks through how throttling decisions ripple through scraper design.

Pro Tip: Log every robots.txt fetch failure with a timestamp and status code. Patterns in those logs often reveal a site's protection changes before your scraper's error rate does.

Expert tips and proxy considerations

Even a scraper that parses robots.txt perfectly still runs into detection systems that have nothing to do with the protocol. Datacenter IP ranges are flagged by many sites regardless of what robots.txt allows, which is where mobile carrier IPs make a practical difference. Some proxy providers route requests through real carrier IPs rather than datacenter ranges, which matters most when you need geo-accurate results from specific US cities rather than a generic exit node. That combination of RFC-compliant parsing on the code side and carrier-grade IPs on the network side is what keeps large scraping jobs from stalling out on both fronts. For teams running at scale, our guide on bot detection signals scrapers must handle covers the patterns that trigger blocks beyond robots.txt rules.

Author perspective: prioritizing compliance, scale, and data quality

Robots.txt compliance is the easy part. The harder call is knowing when automated collection should stop and a human conversation, or a data license, should start. Ignoring robots.txt entirely tends to cost you in IP blocks and reputation long before it costs you in court, but that is not the same as legal safety. Clearer licensing norms for AI training data would help everyone building on scraped content, and until that exists, the safer path is caution over volume.

— Jon

Sources

For the protocol itself, RFC 9309 is the canonical source on syntax, matching rules, and error handling. Google's crawling and robots.txt guidance explains how its crawlers apply the spec in practice, including indexing behavior for blocked pages. MDN's implementation guide covers why robots.txt should never stand in for real access controls. For the legal side, the U.S. Copyright Office's AI initiative is the most current source tracking how scraping for AI training is being evaluated.

FAQ

Is robots.txt illegal to ignore?

No law requires following robots.txt directly, since RFC 9309 standardizes it as a voluntary protocol rather than a binding rule. Ignoring it can still expose you to other legal risks depending on what data you access and how, so voluntary does not mean risk-free.

Is AI scraping illegal?

The legality of large-scale scraping for AI training data is not settled in the United States, and the U.S. Copyright Office is actively studying the issue as part of its ongoing AI initiative. Outcomes depend heavily on what content is collected, how it is used, and which jurisdiction applies.

Is web scraping illegal?

Web scraping itself is not inherently illegal, but its legality depends on what you collect, the site's terms, and applicable data protection or copyright law in your jurisdiction. Respecting robots.txt and rate limits reduces risk but does not by itself make every scraping activity lawful.

Is robots.txt still used?

Yes, robots.txt remains widely used, including by sites that specifically disallow AI training crawlers. Compliance varies in practice, since some crawlers ignore the file or alter their user-agent strings to avoid being blocked by name.

How does robots.txt handle server errors during a fetch?

Under RFC 9309, a 4xx response means the file is unavailable and a crawler may treat the site as unrestricted, while a 5xx response means it is unreachable and the crawler must assume complete disallow. This distinction is why scrapers need separate handling logic for each error category rather than treating all failures the same way.

Recommended