How to measure AI crawlers, referrals and transactions

AI-related activity does not appear in one clean analytics channel. Combine request logs, identifiable referrals and order records, while keeping crawls, customer visits and completed actions separate.

Reviewed October 1, 2026 · 7 min read

On this page7 sections
  1. The three layers, and what sees each one
  2. Layer one: reading crawler traffic in your logs
  3. Which bots are which (and why the distinction matters)
  4. Layer two: AI referral traffic
  5. Layer three: agent transactions
  6. Use platform reports for the questions they answer
  7. What to do with the numbers

Many automated fetches never execute your analytics script, and analytics systems may filter bot activity. Browser agents can behave differently and may run scripts. Use CDN and server logs to inspect requests rather than interpreting an empty analytics segment as proof that no agents visited.

Measure three things separately: automated requests, identifiable visits from assistant links and transactions created through a known integration. They answer different questions and cannot be added together as if each represented a customer.

This guide shows where to collect each signal, how to classify it and how to avoid drawing conclusions the data cannot support.

The three layers, and what sees each one#

LayerWhat it isWhere it shows upWhat misses it
Automated requestsCrawlers and other clients fetching resourcesCDN, WAF and origin logsBrowser analytics may miss or filter them
Referral visitsVisitors arriving through identifiable assistant linksReferrers, campaign data and sessionsSome visits have no usable attribution
Integrated transactionsOrders or bookings through a known client flowBackend events and order recordsA browser session may be absent or incomplete

An API-driven order may have no website session, while a browser-assisted purchase may have one. Record origin in the backend where it is known, then reconcile it with analytics rather than assuming every protocol transaction is invisible to browser measurement.

Layer one: reading crawler traffic in your logs#

Use edge logs when a CDN serves or blocks requests before they reach the origin. The commands below are a quick diagnostic for a traditional combined access-log format. They classify claimed user-agents only; adapt the field parsing to your actual log schema.

# Claimed bot labels in a combined access log, not verified identities
rg -oi 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-User|Claude-SearchBot|PerplexityBot|Perplexity-User' access.log | sort | uniq -c | sort -rn

# Requested paths for one claimed caller (combined-log field layout)
rg -i 'GPTBot' access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -20

# Response codes for selected claimed callers
rg -i 'GPTBot|ClaudeBot|PerplexityBot' access.log | awk '{print $9}' | sort | uniq -c

Investigate repeated failures on important pages. A 403 may be an intended block, a 404 may be a bad URL, and a 200 may contain a challenge rather than useful content. Check the response body, rule and caller identity before diagnosing lost visibility.

Which bots are which (and why the distinction matters)#

Search, training and user-triggered fetching have different purposes. Keep them separate in reports so a training crawl does not get counted as a purchase-intent visit. The crawler policy guide links the provider documentation.

NameOperatorDocumented roleReporting caution
GPTBotOpenAITraining crawlNot a customer visit
OAI-SearchBotOpenAISearch crawlingA fetch does not prove citation
ChatGPT-UserOpenAICertain user-triggered fetchesNot proof of a purchase
ClaudeBotAnthropicTraining crawlKeep separate from search
Claude-SearchBot / Claude-UserAnthropicSearch / user-triggered retrievalSeparate the two where possible
PerplexityBot / Perplexity-UserPerplexitySearch / user-triggered retrievalVerify caller identity and purpose
Google-ExtendedGoogleRobots control tokenNot a separate HTTP user-agent to count
CCBotCommon CrawlDataset crawlingNot direct evidence of an assistant recommendation

Start with your own baseline: requests by verified caller, important URLs reached, response status and time period. Report whether the logs are complete, sampled or retained for only a short window.

Do not compare unlike denominators. Total requests, unique pages, referred sessions and orders measure different activity. A rise in one can reflect a crawler configuration change rather than growing customer demand.

Keep the baseline long enough to see normal variation. Annotate deployments, access-rule changes and campaigns so you can investigate a shift without automatically crediting the latest content edit.

User agents can be forged

A user-agent string is easy to copy. Verify the caller using its documented method, such as published ranges, provider-specific DNS checks or supported signatures. Keep unverified claims in a separate bucket. The verification guide explains why identity and authorization are different.

Layer two: AI referral traffic#

Segment identifiable assistant referrers, including the domains relevant to your audience. The example below uses exact host matching for a small list. It will not identify direct visits, stripped referrers or every assistant product.

const AI_SOURCES = {
  'chatgpt.com': 'ChatGPT',
  'chat.openai.com': 'ChatGPT',
  'perplexity.ai': 'Perplexity',
  'claude.ai': 'Claude',
  'gemini.google.com': 'Gemini',
  'copilot.microsoft.com': 'Copilot',
}

function aiChannel(referrer) {
  if (!referrer) return null
  try {
    const host = new URL(referrer).hostname.replace(/^www\./, '')
    return Object.hasOwn(AI_SOURCES, host) ? AI_SOURCES[host] : null
  } catch {
    return null
  }
}

Treat this as identifiable referral traffic, not a complete AI-influence total. Someone may read a recommendation and later arrive directly. Compare conversion and enquiry quality with suitable baselines, and disclose small samples rather than applying an industry average to your own business.

Layer three: agent transactions#

Record a known integration origin when creating an order or booking. Distinguish the protocol, authenticated client and entry channel where those facts are available. Do not trust a client-supplied 'AI order' flag as verified attribution, and do not store unnecessary personal or payment data just for reporting.

Use stable order identifiers to reconcile completed actions, cancellations and refunds without double counting retries. Some past attribution can be reconstructed from retained logs; information you never recorded may remain unknown. Label that uncertainty explicitly.

Check access alongside measurement

A Nexez scan reports how its public requests are handled and what information they extract. Combine that diagnostic with actual caller logs and your own order evidence.

Scan your site free

Use platform reports for the questions they answer#

Search and merchant platforms may provide their own visibility or AI-related reports. Check the definition, account availability, geography and reporting window of the feature you actually have. Do not assume every account can isolate every AI placement.

A platform's share-of-voice metric, your request count and your referral sessions have different denominators. Keep them side by side with their definitions rather than forcing them into a single funnel whose steps cannot be matched.

What to do with the numbers#

Use the evidence to choose an investigation, not to jump straight to a cause:

  1. Access failures: identify the caller, path and matched rule, then decide whether the result is intended.
  2. Coverage gaps: compare requested pages with important URLs, while accounting for log sampling, caches and alternate data sources.
  3. Stale information: inspect the current page, feed and cited source before assuming a recrawl delay.
  4. Low referrals despite requests: consider crawl purpose, answers without clicks and missing attribution, not just content quality.
  5. Weak conversions: inspect the landing page and customer journey; low agent-order counts alone do not prove an API is needed.

The useful report distinguishes what happened from what you infer. Show known requests, identifiable visits and verified outcomes, then list the gaps that prevent stronger conclusions. That makes the next investment easier to assess.

Connect published offers with observable outcomes

See how Nexez publishes structured listings and supported actions. Review the reporting available for the workflows you use and reconcile it with your own business records.

See how it works

Frequently asked questions

Why does Google Analytics show no AI crawler traffic?

Many crawlers do not run the analytics script, and analytics products may filter automated activity. Browser agents can behave differently. Use CDN and server logs to inspect requests; a zero in analytics is not proof that no crawler reached the site.

How do I tell AI crawlers apart in my server logs?

Use user-agent fields as an initial classification, then apply the provider's documented verification method. Separate verified, unverified and unknown requests. Google-Extended is a robots control token, not a separate HTTP crawler to count.

What is the difference between GPTBot and ChatGPT-User?

GPTBot is documented for training, ChatGPT-User for certain user-triggered fetches and OAI-SearchBot for search. Report them separately because their requests do not represent the same intent. A fetch alone does not prove a citation, customer visit or purchase.

Can AI crawler traffic be faked?

Yes. A request can claim a familiar user-agent. Verify using the mechanism documented for that specific caller and retain a separate unverified category. Do not grant privileged access or report confirmed provider activity from the name alone.

How do I track sales that AI agents complete?

Record known integration origin and authenticated client information in the order system, then reconcile completed orders, cancellations and refunds. Some browser-assisted purchases also create sessions. Avoid counting retries twice or trusting an arbitrary client-supplied attribution label.

Is AI crawler traffic large enough to bother measuring?

Measure a small baseline before deciding. The cost, customer value and volume vary by site. A simple report of verified requests, identifiable referrals and known transactions is more useful than assuming a network-wide growth statistic predicts your business.

Keep reading