Which AI crawlers should your business allow?
Search crawlers, training crawlers and user-triggered fetchers serve different purposes. Choose access for each, then check that your robots.txt, CDN and application rules produce the result you intended.
Reviewed October 1, 2026 · 7 min read
On this page7 sections
A public business website often benefits when an assistant can read its services, prices and hours. That does not settle whether you want the same content used for model training or how you should protect private and transactional routes.
Make those decisions separately. Clear access rules help avoid accidental blocks while leaving room for rate limits, authentication and content-use preferences. Our website study illustrates why a scanner's reachability result also needs careful interpretation.
This guide covers the main distinctions, an example training preference and the checks needed after a policy change.
Separate three kinds of activity#
Provider names alone are not enough. One company may operate several callers with different documented purposes:
| Category | What it does | What blocking costs you | Sensible default |
|---|---|---|---|
| Search | Retrieves information for search experiences | Can reduce that provider's access to your pages | Allow public discovery where it serves your goals |
| User-triggered fetch or agent | Retrieves or operates a page for a user | Can interrupt a customer workflow | Allow appropriate public paths with suitable controls |
| Training | Collects data for model development | Expresses or enforces a training-access preference | Decide separately from search |
Blocking a training-only crawler is not the same as withdrawing from search. Conversely, allowing a search crawler is not permission for an agent to make purchases or access customer data.
The roster, by who sends it#
Check the providers' own documentation: OpenAI, Anthropic and Perplexity. The following names are useful policy references, but not all are separate HTTP user-agents.
| User agent | Operator | Job | Obeys robots.txt |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Search: powers ChatGPT search answers | Yes |
| ChatGPT-User | OpenAI | Agent: fetches a page a user just asked about | May not apply |
| GPTBot | OpenAI | Training | Yes |
| Claude-SearchBot | Anthropic | Search: improves Claude search results | Yes |
| Claude-User | Anthropic | Agent: retrieves pages for a user’s question | Yes |
| ClaudeBot | Anthropic | Training | Yes |
| PerplexityBot | Perplexity | Search: surfaces and links your site | Yes |
| Perplexity-User | Perplexity | Agent: user-initiated visit | Generally ignores it |
| Googlebot | Search, and everything Google builds on Search | Yes | |
| Google-Extended | Robots control token for specified Gemini uses, not a separate HTTP crawler | A robots.txt control token | |
| Storebot-Google | Shopping and the Shopping tab | Yes |
Read the role as well as the name. A familiar user-agent is a claim made by the request, not authenticated proof of the operator.
OpenAI and Perplexity distinguish user-triggered fetching from background crawling and document different robots behavior for those requests. A request being user-triggered does not make it an ordinary human browser or exempt it from your access controls.
Google-Extended does not do what its name suggests
Google-Extended is a robots control for specified Gemini uses, separate from Googlebot's Search role. It is not an AI Mode opt-out or a separate user-agent to count in logs. Search indexing and preview controls have their own effects; use Google's current guidance for the outcome you need.
Account for Cloudflare's updated controls#
Cloudflare's September rollout introduced Disallow AI Training as distinct from Block. Its new behavior should not be inferred solely from the July announcement. The updated Cloudflare guide explains the distinction and verification steps.
A full network block can affect a crawler that serves search as well as training. A content-use preference can preserve access while expressing a narrower choice. Check which effect your selected control actually has.
Also inspect custom WAF rules and bot challenges. A permissive category setting may not override a separate rule that denies the same request.
Review the effective rules, not just one toggle
Record the settings for each relevant domain, the published robots.txt and the matched rules in recent security events. Ask your host or agency for those details if they manage the edge.
An example training-only robots preference#
The example below asks selected training callers not to crawl. Merge it carefully with your existing file. It does not grant additional access to search, protect private pages or override CDN rules. Removing these groups would restore whatever policy otherwise applies, not keep them blocked.
# Example only: merge with the site's existing groups. # These training preferences are separate from search access. User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended Disallow: / # Preserve the site's existing general and search-specific rules. Sitemap: https://example.com/sitemap.xml
If you want to allow training, use the policy that reflects that choice rather than copying the example unchanged. Specific user-agent groups can change which rules a crawler applies, so review path restrictions when adding one.
Serve the appropriate robots.txt for each origin. Keep intended public resources reachable, but do not use robots.txt to protect private information: authentication and authorization belong in the application and access layer.
Test a public fetch and inspect the result
A Nexez scan reports what its requests can reach and extract. Pair that result with real provider logs and search tools to understand how different callers encounter the site.
Scan your site freeWhen blocking is the right answer#
Blocking or limiting access can be appropriate for private content, excessive load or a content-use policy. Choose the scope and mechanism deliberately:
- Use authentication and authorization for private or customer-specific content.
- Use documented training controls to express training preferences where supported.
- Use rate limits and capacity controls for excessive request load.
- Evaluate any licensing or paid-access arrangement against its actual terms and demand.
Recheck public customer journeys after a change. A policy can be intentional and still broader than expected if it catches search, user-triggered requests or shared infrastructure.
robots.txt is a request, not a lock#
Robots.txt is a published instruction for compliant consumers, not a lock. Provider behavior differs, especially for user-triggered requests. It cannot establish a caller's identity or prevent a noncompliant client from sending a request.
Use the verification method documented for the particular caller: published address ranges, provider-specific DNS validation or supported request signatures. Do not invent one universal test for all bots, and do not allow privileged access based only on a user-agent string.
Signed requests can help verify identity where supported, but they do not prove a customer's authorization for every action. The agent verification guide separates those questions.
Then measure it, because the answer changes#
Review the policy after provider changes, a CDN migration or unexplained access errors. Keep an owner and a dated record of the settings you intended.
Use server and CDN evidence for automated fetches. JavaScript analytics often miss them or filter them, while browser agents may execute scripts. Neither a zero nor a spike in analytics alone settles the question.
The measurement guide covers response codes, verification and referral attribution. Start with legitimate requests to important pages and inspect failures rather than counting every claimed bot hit as useful traffic.
A 403 can indicate a deliberate policy or an accidental block; a 200 can contain a challenge or empty page. Read the response and matched rule before changing access.
Once intended access works, review the information delivered: useful text, accurate structured data and a clear next step. Reachability is one prerequisite, not proof of citation or conversion.
Make accessible pages useful to customers
See how Nexez publishes structured business listings and supported actions. Pair clear information with verified access and a tested customer flow.
See how it worksFrequently asked questions
Should a small business block AI crawlers?
Separate the decisions. Public search and user-triggered access may help customers find and evaluate you, while training use is a different preference. Protect private routes and control excessive load regardless of the caller's brand name.
Does blocking Google-Extended remove me from AI Overviews?
No. Google-Extended controls specified Gemini uses and is separate from Search crawling. Use Google's current indexing and preview controls for Search outcomes. Do not assume that blocking one token removes all AI uses or that leaving Search is the only possible control.
What did Cloudflare change in September 2026?
It introduced updated granular controls, including Disallow AI Training, which differs from a full block. Inspect your migrated or selected settings and other edge rules. The Cloudflare guide explains why older summaries can be misleading.
What is the difference between GPTBot and ChatGPT-User?
GPTBot is documented for training, while ChatGPT-User handles certain user-triggered fetches. OAI-SearchBot has a separate search role. These distinctions let you choose different policies, subject to each caller's documented behavior and your own access controls.
Why do my logs show AI crawlers getting 403 errors?
Inspect the matched CDN, WAF or application rule. A 403 may reflect bot protection, authentication, a custom deny rule or another condition. Robots.txt alone does not generate a standard access denial. Verify the caller and response before changing the policy.
Can I get paid instead of blocking?
Some providers offer licensing or paid-access products. Availability and economics depend on the content, traffic and program terms. Evaluate them separately from ordinary search access; blocking requests does not automatically create paying demand.
How often should I revisit this?
Review after relevant provider or infrastructure changes and when monitoring shows unexplained failures. A scheduled review can help, but real request errors and renamed or newly documented callers should trigger an earlier check.
Keep reading
AI readiness study: what we found on 652 small-business websites
We checked 652 small-business websites across five industries for public access, structured information and selected agent-discovery files.
9 min readAgent readinessCloudflare's September 15 changes: what to check now
Cloudflare's September rollout changed the choices for search, training and agent traffic.
6 min readAgent readinessHow to verify an AI agent is who it says it is
Check who sent an agent request, then decide what that caller is allowed to do. A verified identity is only the first step.
8 min read