Which AI crawlers should your business allow?

Search crawlers, training crawlers and user-triggered fetchers serve different purposes. Choose access for each, then check that your robots.txt, CDN and application rules produce the result you intended.

Reviewed October 1, 2026 · 7 min read

On this page7 sections
  1. Separate three kinds of activity
  2. The roster, by who sends it
  3. Account for Cloudflare's updated controls
  4. An example training-only robots preference
  5. When blocking is the right answer
  6. robots.txt is a request, not a lock
  7. Then measure it, because the answer changes

A public business website often benefits when an assistant can read its services, prices and hours. That does not settle whether you want the same content used for model training or how you should protect private and transactional routes.

Make those decisions separately. Clear access rules help avoid accidental blocks while leaving room for rate limits, authentication and content-use preferences. Our website study illustrates why a scanner's reachability result also needs careful interpretation.

This guide covers the main distinctions, an example training preference and the checks needed after a policy change.

Separate three kinds of activity#

Provider names alone are not enough. One company may operate several callers with different documented purposes:

CategoryWhat it doesWhat blocking costs youSensible default
SearchRetrieves information for search experiencesCan reduce that provider's access to your pagesAllow public discovery where it serves your goals
User-triggered fetch or agentRetrieves or operates a page for a userCan interrupt a customer workflowAllow appropriate public paths with suitable controls
TrainingCollects data for model developmentExpresses or enforces a training-access preferenceDecide separately from search

Blocking a training-only crawler is not the same as withdrawing from search. Conversely, allowing a search crawler is not permission for an agent to make purchases or access customer data.

The roster, by who sends it#

Check the providers' own documentation: OpenAI, Anthropic and Perplexity. The following names are useful policy references, but not all are separate HTTP user-agents.

User agentOperatorJobObeys robots.txt
OAI-SearchBotOpenAISearch: powers ChatGPT search answersYes
ChatGPT-UserOpenAIAgent: fetches a page a user just asked aboutMay not apply
GPTBotOpenAITrainingYes
Claude-SearchBotAnthropicSearch: improves Claude search resultsYes
Claude-UserAnthropicAgent: retrieves pages for a user’s questionYes
ClaudeBotAnthropicTrainingYes
PerplexityBotPerplexitySearch: surfaces and links your siteYes
Perplexity-UserPerplexityAgent: user-initiated visitGenerally ignores it
GooglebotGoogleSearch, and everything Google builds on SearchYes
Google-ExtendedGoogleRobots control token for specified Gemini uses, not a separate HTTP crawlerA robots.txt control token
Storebot-GoogleGoogleShopping and the Shopping tabYes

Read the role as well as the name. A familiar user-agent is a claim made by the request, not authenticated proof of the operator.

OpenAI and Perplexity distinguish user-triggered fetching from background crawling and document different robots behavior for those requests. A request being user-triggered does not make it an ordinary human browser or exempt it from your access controls.

Google-Extended does not do what its name suggests

Google-Extended is a robots control for specified Gemini uses, separate from Googlebot's Search role. It is not an AI Mode opt-out or a separate user-agent to count in logs. Search indexing and preview controls have their own effects; use Google's current guidance for the outcome you need.

Account for Cloudflare's updated controls#

Cloudflare's September rollout introduced Disallow AI Training as distinct from Block. Its new behavior should not be inferred solely from the July announcement. The updated Cloudflare guide explains the distinction and verification steps.

A full network block can affect a crawler that serves search as well as training. A content-use preference can preserve access while expressing a narrower choice. Check which effect your selected control actually has.

Also inspect custom WAF rules and bot challenges. A permissive category setting may not override a separate rule that denies the same request.

Review the effective rules, not just one toggle

Record the settings for each relevant domain, the published robots.txt and the matched rules in recent security events. Ask your host or agency for those details if they manage the edge.

An example training-only robots preference#

The example below asks selected training callers not to crawl. Merge it carefully with your existing file. It does not grant additional access to search, protect private pages or override CDN rules. Removing these groups would restore whatever policy otherwise applies, not keep them blocked.

# Example only: merge with the site's existing groups.
# These training preferences are separate from search access.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /

# Preserve the site's existing general and search-specific rules.
Sitemap: https://example.com/sitemap.xml

If you want to allow training, use the policy that reflects that choice rather than copying the example unchanged. Specific user-agent groups can change which rules a crawler applies, so review path restrictions when adding one.

Serve the appropriate robots.txt for each origin. Keep intended public resources reachable, but do not use robots.txt to protect private information: authentication and authorization belong in the application and access layer.

Test a public fetch and inspect the result

A Nexez scan reports what its requests can reach and extract. Pair that result with real provider logs and search tools to understand how different callers encounter the site.

Scan your site free

When blocking is the right answer#

Blocking or limiting access can be appropriate for private content, excessive load or a content-use policy. Choose the scope and mechanism deliberately:

  • Use authentication and authorization for private or customer-specific content.
  • Use documented training controls to express training preferences where supported.
  • Use rate limits and capacity controls for excessive request load.
  • Evaluate any licensing or paid-access arrangement against its actual terms and demand.

Recheck public customer journeys after a change. A policy can be intentional and still broader than expected if it catches search, user-triggered requests or shared infrastructure.

robots.txt is a request, not a lock#

Robots.txt is a published instruction for compliant consumers, not a lock. Provider behavior differs, especially for user-triggered requests. It cannot establish a caller's identity or prevent a noncompliant client from sending a request.

Use the verification method documented for the particular caller: published address ranges, provider-specific DNS validation or supported request signatures. Do not invent one universal test for all bots, and do not allow privileged access based only on a user-agent string.

Signed requests can help verify identity where supported, but they do not prove a customer's authorization for every action. The agent verification guide separates those questions.

Then measure it, because the answer changes#

Review the policy after provider changes, a CDN migration or unexplained access errors. Keep an owner and a dated record of the settings you intended.

Use server and CDN evidence for automated fetches. JavaScript analytics often miss them or filter them, while browser agents may execute scripts. Neither a zero nor a spike in analytics alone settles the question.

The measurement guide covers response codes, verification and referral attribution. Start with legitimate requests to important pages and inspect failures rather than counting every claimed bot hit as useful traffic.

A 403 can indicate a deliberate policy or an accidental block; a 200 can contain a challenge or empty page. Read the response and matched rule before changing access.

Once intended access works, review the information delivered: useful text, accurate structured data and a clear next step. Reachability is one prerequisite, not proof of citation or conversion.

Make accessible pages useful to customers

See how Nexez publishes structured business listings and supported actions. Pair clear information with verified access and a tested customer flow.

See how it works

Frequently asked questions

Should a small business block AI crawlers?

Separate the decisions. Public search and user-triggered access may help customers find and evaluate you, while training use is a different preference. Protect private routes and control excessive load regardless of the caller's brand name.

Does blocking Google-Extended remove me from AI Overviews?

No. Google-Extended controls specified Gemini uses and is separate from Search crawling. Use Google's current indexing and preview controls for Search outcomes. Do not assume that blocking one token removes all AI uses or that leaving Search is the only possible control.

What did Cloudflare change in September 2026?

It introduced updated granular controls, including Disallow AI Training, which differs from a full block. Inspect your migrated or selected settings and other edge rules. The Cloudflare guide explains why older summaries can be misleading.

What is the difference between GPTBot and ChatGPT-User?

GPTBot is documented for training, while ChatGPT-User handles certain user-triggered fetches. OAI-SearchBot has a separate search role. These distinctions let you choose different policies, subject to each caller's documented behavior and your own access controls.

Why do my logs show AI crawlers getting 403 errors?

Inspect the matched CDN, WAF or application rule. A 403 may reflect bot protection, authentication, a custom deny rule or another condition. Robots.txt alone does not generate a standard access denial. Verify the caller and response before changing the policy.

Can I get paid instead of blocking?

Some providers offer licensing or paid-access products. Availability and economics depend on the content, traffic and program terms. Evaluate them separately from ordinary search access; blocking requests does not automatically create paying demand.

How often should I revisit this?

Review after relevant provider or infrastructure changes and when monitoring shows unexplained failures. A scheduled review can help, but real request errors and renamed or newly documented callers should trigger an earlier check.

Keep reading