APIs, integration & security — in depth

Robots.txt and Crawl Policy Compliance in Agentic Systems

AI crawlers now ignore the honor system that made robots.txt work for search engines.

Reporter · · 12 min read
Cover illustration for “Robots.txt and Crawl Policy Compliance in Agentic Systems”
Live Web Reasoning · September 30, 2026 · 12 min read · 2,623 words

Robots.txt was built in 1994 as a polite request. It asked search engine crawlers to stay out of certain folders, and it became a standard, RFC 9309, without ever picking up legal force or a technical enforcement mechanism of its own. A crawler that wants to ignore the file can simply ignore it. Nothing in the protocol stops that.

That gap didn't matter much for three decades because the crawlers reading robots.txt were almost all search engines, and search engines had a reason to comply: they needed the site owner's goodwill to keep indexing new pages. The relationship was reciprocal. A crawler indexed a page, a user found that page in search results, and the site got a visitor.

That reciprocity is what's broken now. Many AI crawlers ingest a page's content to produce an answer somewhere else, an answer that never sends a visitor back to where the content came from. The crawl-for-traffic bargain that made robots.txt workable on the honor system doesn't apply to a bot whose whole purpose is to absorb text into a model and never link back.

So the file is still doing what it did in 1994, issuing a polite, unenforceable request, but the population making requests against it now includes actors with no dependency on staying in a site owner's good graces. For a site owner, that's a housekeeping problem: which bots to allow, which to block. For anyone building a system that sends automated agents out to fetch live pages, it's an architectural problem. The agent's fetch behavior encodes a policy stance whether the developer wrote one or not, and figuring out what stance to encode requires understanding a crawler landscape that has splintered into pieces robots.txt was never built to distinguish.

The crawler landscape's purpose-based split

The crawler population used to be one thing: search engines, indexing pages, mostly Google. It's now split along a line that matters enormously for anyone writing a robots.txt file or designing a fetcher: purpose. The biggest operators run separate bots for training a model and for retrieving pages to answer a live query, and each bot carries its own user-agent string that a site can allow or block independently.

OpenAI runs GPTBot for model training and OAI-SearchBot for ChatGPT's search citations, two different bots with two different jobs. Anthropic runs three: ClaudeBot for training, Claude-SearchBot for search indexing, and Claude-User for the case where a person asks Claude to go read a specific page right now. Amazon splits Amazonbot from Amzn-SearchBot. Google's Google-Extended token controls opt-out for Gemini training and grounding specifically, and notably does not cover AI Overviews or AI Mode, which run through the separate Googlebot pathway entirely.

A third category exists: the real-time, user-triggered fetch. When someone asks an AI assistant to summarize a specific URL, that request looks a lot like a person clicking a link, not a crawler harvesting a dataset. Providers often treat it as user-directed access rather than crawling, so the robots.txt rules a site owner wrote for "crawlers" may not apply to it.

The stakes of getting this taxonomy wrong are already visible. Search Engine Journal's reporting, cited in the Decision Matrix analysis, found that a majority of top news publishers block at least one retrieval or search bot while intending to block only training bots, which quietly drops them out of AI-powered answer engines without anyone deciding that should happen. Cloudflare's network research adds an economic dimension: Anthropic's training crawler generates a crawl-to-referral ratio orders of magnitude worse than Googlebot's, pulling enormous volumes of content while sending almost nothing back in the form of visits. That's a concrete argument for blocking training specifically, not AI broadly.

Common Crawl's bot, CCBot, deserves its own line here because its effects are delayed and easy to miss. Blocking it removes a site from future dumps of the Common Crawl dataset, which feeds a large share of open-source language models, but anything already archived stays archived. The site owner blocking it today won't see that consequence in their own logs.

Put together, this is the map a developer needs before making any single robots.txt or fetch-policy decision: training bots, search bots, and user-triggered fetches, each behaving differently and each carrying different consequences for both the fetching and the fetched.

The wrong default: blanket "block all AI bots" rules for agentic pipelines

The "block all AI bots" instinct made sense back when there was exactly one category of AI crawler to worry about. Applied to the landscape described above, it conflates two decisions that have opposite consequences. Blocking GPTBot keeps a site's content out of OpenAI's training data. Blocking OAI-SearchBot in the same breath removes the site from ChatGPT's search answers entirely. The same split happens with Claude-SearchBot and Claude's web answers, and with PerplexityBot and Perplexity's citations, and none of these are training bots, yet all are caught by a blanket Disallow.

Developers building retrieval-augmented generation pipelines or any agentic system that fetches live pages feel this from both sides. The agents developers build will encounter these same blocks when fetching sources, and the products they serve may become invisible to competitor agents that surface answers to users. Block policy affects what a developer's product can see, and it affects whether a developer's product gets seen.

The user-triggered fetch category makes this messier still. An agent fetching a URL because a user explicitly asked it to may not be covered by the robots.txt logic a developer wrote for "crawler" behavior at all, since providers handle that case differently and there's no guarantee the target site's policy will even be read in that context.

The defensible default for 2026 threads this needle by using the same purpose split that created the problem: block the training bots (GPTBot, ClaudeBot, CCBot, Google-Extended, Applebot-Extended) while allowing the search and retrieval bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot) to keep working. Robots.txt can express that distinction precisely now, because the underlying user-agents are separated at the source. The rule just has to be written with that separation in mind instead of reaching for a single blanket line.

The emerging policy stack above and below robots.txt

Robots.txt handles access: fetch this path, don't fetch that one. It has no vocabulary for permission at the level AI use cases actually need, questions like whether a page can be used for training versus real-time retrieval versus nothing at all. Two newer files have shown up specifically to fill that gap.

ai.txt is a plain-text file, placed either at the site root under a 2023 proposal from Spawning or at /.well-known/ai.txt under a 2026 IETF draft, that carries tags like No-Training, No-Inference, and Allow-RAG. Those tags give a site owner purpose-based permission controls that robots.txt's path-based Disallow simply cannot express. Its legal weight is also different in kind: ai.txt's status is tied to the EU AI Act and to Text and Data Mining provisions, while robots.txt remains a norm everyone follows by convention, with no statute behind it.

llms.txt does something else entirely, and conflating it with ai.txt is a common mistake. Proposed by Jeremy Howard as a Markdown file at the site root, it controls nothing. It's a navigation aid, pointing AI systems toward a site's cleanest, highest-value content at the moment of inference. About one in ten domains have adopted it so far, but publishing one is low-risk and genuinely useful for the narrow purpose of helping an inference-time agent find good content instead of scraping a full rendered page and hoping for the best.

The layer below robots.txt actually enforces something. A CDN or WAF rule gets evaluated before robots.txt is ever read, so a Cloudflare block overrides a robots.txt Allow regardless of what the text file says, and for crawlers that ignore robots.txt outright, IP and WAF blocking is the only defense that actually works. Cloudflare's own three-way classification, splitting traffic into Search, Agent, and Training categories, lets a site manage each independently, and starting September 15, 2026, newly onboarded Cloudflare domains and all existing free-tier customers block both Agent and Training traffic by default on any page that shows ads.

Laid out by enforcement strength, the stack runs: WAF and CDN rules first, enforced before anything else loads; then server-level rate limits and IP blocks; then robots.txt, honored by whichever bots choose to honor it; then ai.txt, honored by whoever has signed on to the emerging standard; then llms.txt, which never enforced anything to begin with. A developer relying solely on robots.txt compliance is relying on the weakest layer in that entire stack.

Diagram: The Enforcement Stack: Strongest to Weakest. Visualizes: Show five layers of crawler enforcement arranged by actual enforcement strength, from strongest to weakest.

Since August 2, 2026, AI providers that ignore robots.txt face fines up to €15 million or a percentage of worldwide annual turnover, whichever is higher, and that exposure lands on General-Purpose AI providers, not on every developer running a scraping script.

The connective tissue between that legal obligation and day-to-day development practice is the General-Purpose AI Code of Practice, published July 10, 2025. Its copyright chapter, Measure 1.3, commits signatories to deploy crawlers that actually read and follow robots.txt. OpenAI, Anthropic, Google, Microsoft, Amazon, and IBM signed it. Meta refused publicly.

The obligations run deeper than just reading the file. Under Article 53 of the EU AI Act, GPAI providers have to publish a detailed summary of what content trained their models, following a template the AI Office provides, and separately maintain a copyright-compliance policy covering text-and-data-mining opt-outs. A compliant training set should carry a metadata record for every document: source URL, crawl timestamp, robots_txt_status, a hash of the robots.txt snapshot at crawl time, detected license, PII removal status, and TDMRep opt-out flag, under the Article 53 framework.

GDPR runs alongside this as a separate constraint. European data-protection guidance is clear that a page being publicly accessible doesn't exempt an operator from GDPR obligations once that page contains personal data; scraping something public doesn't make training on the personal information inside it lawful by default.

The US picture is far less settled. NYT v. OpenAI and Microsoft, along with several other pending lawsuits, will shape how fair use gets applied to AI training once they resolve. States are drafting their own legislation in the meantime, and the Copyright Office has put out guidance and policy reports addressing AI training and fair use, without anything close to a settled framework yet.

None of this touches a small RAG pipeline directly. But it changes the calculation for anyone feeding data into a GPAI system or building toward that scale: robots.txt compliance moves from a courtesy to a documented legal input, one a regulator can pull into an audit.

Compliance's structural weakness: user-agent spoofing and verification

Every control described so far, robots.txt, ai.txt, even the EU's fine structure, rests on one identity mechanism: the user-agent string a client sends with its request. Nothing in HTTP binds that string to who actually operates the client. Any crawler can claim to be any bot, and there's no cryptographic check stopping it.

That unauthenticated foundation is visible in the actual numbers. A 2026 analysis cited by openhermit.com found up to 72% of AI crawlers violate robots.txt rules. Compliance is a property of the individual vendor's implementation, not something guaranteed by category membership. Behavior varies even across products from the same company.

The iFixit case from 2024 shows what that looks like on the ground. iFixit's CEO reported ClaudeBot hitting the company's site close to a million times in a single day, which he called a breach of the site's terms of service. Freelancer.com and Read the Docs reported similar surges around the same time, all from a bot whose official documentation says it respects robots.txt.

That documentation-versus-behavior gap becomes a legal problem once the EU AI Act's enforcement mechanism is in play. Any dispute over whether a provider actually respected robots.txt turns into an argument about which log entries belong to which operator, because the user-agent header is the only attribution signal on the table, and it's entirely self-reported.

Enforcement layers don't even agree with each other reliably. Cloudflare's network research, cited by dataimpulse.com, found a meaningful share of sites accidentally blocking major AI crawlers at the CDN layer while their robots.txt file still says Allow. The two layers often don't line up.

For a developer building an agentic pipeline, this cuts both directions at once. There's no reliable way to verify that a source site's robots.txt is being honored by other crawlers touching the same data, and there's no guarantee that a developer's own compliance signals get attributed correctly if the pipeline runs through a managed API or proxy layer that rotates identifiers behind the scenes. The mechanism being leaky is not a reason to give up on compliance. Systems need to document their own compliance posture explicitly, because the identity layer underneath cannot be trusted to do that documenting on its own.

Decisions agentic pipeline developers must make for live web data fetches

Every fetch an agent makes carries a policy stance baked into it, whether a developer chose that stance deliberately or just inherited whatever the HTTP library does by default. Given everything above, that inheritance is a bad default. Five decisions need to be made explicitly rather than left to chance.

The first is whether to honor robots.txt directly or delegate that responsibility. A pipeline can respect robots.txt for every path it touches, crawl-delay directives included, or it can hand that job to a managed API that handles compliance as a service. Either is defensible, but leaving it unexamined means inheriting someone else's defaults without knowing what they are, which carries legal and reputational consequences a developer should choose knowingly.

The second is separating training intent from retrieval intent. Content pulled for model training and content pulled for real-time RAG grounding can carry different obligations under the same law. If a pipeline feeds a GPAI system, Article 53's documentation requirement kicks in, and the pipeline needs to emit a metadata record at fetch time containing source URL, crawl timestamp, robots_txt_status, a snapshot hash of the robots.txt file, detected license, PII removal status, and TDMRep opt-out flag.

The third concerns what happens when the agent actually hits a blocked URL during a run. A well-built pipeline doesn't silently skip it. It surfaces the block as a structured event, logs the policy signal, retries with the correct delay if crawl-delay is specified, or escalates to a human or a different data source if access is flatly denied.

The fourth covers the newer signals. If a target site publishes an ai.txt file with No-Inference or No-RAG tags, the pipeline's response to those tags needs to be written down and followed, not quietly ignored because the format is unfamiliar. llms.txt works differently: it's a navigation aid a pipeline can use actively, pointing it toward cleaner content and cutting down on failed scrapes of messy rendered pages.

The fifth is distinguishing the reason behind a failed fetch. A 4xx response can mean a CDN-enforced block that no robots.txt compliance will resolve, a robots.txt disallow representing a policy choice the operator has made, or a rate limit that calls for backoff and retry. Treating all three the same way, usually by giving up, throws away information the pipeline needs to behave correctly next time.

A customer support retrieval system built on this logic treats RAG not as a fixed step in a pipeline but as a tool the model can choose to invoke inside its own reasoning loop, deciding when to retrieve and how to query for it. That design choice puts the compliance logic at the tool layer itself, where the fetch actually happens. At that scale, robots.txt compliance is no longer a courtesy extended to site owners but a property of the system's own architecture.

Sources

  1. Beyond Robots.txt: Implementing AI.txt and LLMs.txt for Purpose-Based Scraping Control
  2. Robots.txt, AI Crawlers & Web Scraping in 2026
  3. AI Crawler Access Control: The 2026 Decision Matrix
  4. Robots.txt for AI Agents 2026: The Agent-Allow Strategy — OpenHermit Blog

More in Live Web Reasoning