Technical October 4, 2026 · 15 min read

AI Crawler Log Analysis: See Which Bots Hit Your Store

GA4 cannot see an AI crawler, because crawlers do not run JavaScript. Your logs can. Here is how to read Cloudflare's AI Crawl Control and your raw server logs to find out which assistants are reading your store, what they read, whether they obey robots.txt, and which ones are worth the bandwidth.

3 Crawler Categories: Training, Search, Agent
0 Crawlers Visible In GA4
24h Metrics Window On Cloudflare's Free Plan
1 Ratio That Decides Allow Or Block
Quick Answer

AI crawlers never appear in GA4 because they do not execute JavaScript, so the only way to know which assistants read your store is the request log. If your domain is on Cloudflare, AI Crawl Control (free on every plan as of September 2026) shows which AI crawlers hit you, how often, which pages, and whether they respected robots.txt, with the Crawlers, Metrics and Robots.txt tabs; the Free plan limits metrics to the last 24 hours and detects crawlers by user agent only. If you are not on Cloudflare, or want history, pull the raw access log from your host and count requests by user agent with a few shell commands. Then sort every crawler into one of three categories: training (feeds a model; produces no referral), search or retrieval (fetches pages to answer questions and can cite you), and agent (fetches a page on a user's behalf). Compute the crawl-to-referral ratio for each retrieval bot against GA4's AI referrals, keep the ones that send visitors, and decide the training bots on principle rather than analytics, because they will never send anyone.

An assistant that cites you had to read you first. The reading is in your logs, dated, by page, and most stores have never looked.

The AI crawler list on this site names the bots and what each one does. This is the follow-up it promised: how to see them on your own store. It uses Cloudflare because that is what we run and what most Shopify and WordPress stores in our client base run, and it covers raw logs for everyone else. Nothing here needs a developer; it needs a dashboard login and, for the raw-log path, a terminal and ten minutes.

Disclosure And Dates

Evolve Media Agency sells AI search visibility work, which includes exactly this kind of audit. Cloudflare's features and interface labels are described as documented on September 15, 2026, and Cloudflare changes them often; the docs linked in the sources are current where this post is not. We have no relationship with Cloudflare beyond being a customer.

01/12 Section

Why This Is A Log Question, Not An Analytics One

GA4 fires from a JavaScript tag in the browser. Crawlers request the HTML and leave; they do not run the tag. So the assistant that read forty of your product pages last night left no trace in GA4, and the only record is the web server's or the CDN's request log: timestamp, IP, user agent, path, status code, bytes.

That record answers questions analytics never can. Which assistants know your catalog. How often they refresh it. Which pages they think are worth re-reading. Whether the crawler you blocked in robots.txt kept crawling. And, joined to GA4's referral data, whether the reading ever turns into a visitor, which is the number that decides whether to keep letting it read.

Definition

AI crawler. A bot operated by an AI company that requests pages from your site for one of three purposes: to collect training data for a model, to retrieve current content so an assistant can answer a question and cite the source, or to fetch a page an assistant's user asked it to open. Each identifies itself with a user agent string (GPTBot, ClaudeBot, PerplexityBot and so on), most publish the IP ranges they crawl from, and their behavior toward robots.txt varies by operator and purpose.

02/12 Section

The Three Kinds Of AI Bot

Every decision later in this post depends on which category a crawler is in, and Cloudflare adopted the same three-way split in its July 2026 bot controls.

CategoryWhat It DoesExamples (User Agents)Sends Referrals?Default Stance For A Store
TrainingCollects pages to train or fine-tune a model; nothing you see later is tied to a visitGPTBot, ClaudeBot, Google-Extended (a control token rather than a bot), CCBot, Applebot-Extended, Meta-ExternalAgent, BytespiderNoA principle decision: allow if you want future models to know your brand; block if you object to training use
Search / retrievalFetches current pages to answer a live question and cite the sourceOAI-SearchBot, PerplexityBot, Claude-SearchBot, Bingbot (for Copilot), Googlebot (for AI Overviews and AI Mode)Yes, this is where AI referral traffic comes fromAllow, and measure the ratio
Agent / user-triggeredFetches a specific page because a user asked the assistant to open, summarize or shop itChatGPT-User, Claude-User, Perplexity-User, and browser-agent trafficSometimes; often the user follows throughAllow; blocking this blocks a customer

The operator names shift and new bots appear monthly; the categories do not. When a user agent you do not recognize shows up, the first question is which of the three it is, and the operator's published crawler page usually says. The robots.txt guide has the current token for each and the directive that controls it.

03/12 Section

Cloudflare AI Crawl Control: What Each Tab Shows

AI Crawl Control (the product formerly called AI Audit) is in the Cloudflare dashboard under your domain and, as of September 2026, is available on every plan including Free. It has three views that matter for analysis and one for action.

AI Crawl Control Dashboard, Per Domain
Tab 01
Crawlers

Every known AI crawler that requested your domain, with operator, category, request count and trend, and a per-crawler allow or block control. This is the list you are here for.

Tab 02
Metrics

Volume, bandwidth and, on paid plans, request-to-referral patterns over time. On Free the window is the last 24 hours; paid plans get longer ranges. Screenshot it daily if you are on Free and want history.

Tab 03
Robots.txt

Health of your robots.txt (status code, whether one exists per hostname), whether it carries Content Signals, and, the useful part, which crawlers requested paths your robots.txt disallows, with the path and the count.

Action
Block, Allow, Or Respond

Per-crawler block and allow on every plan; on paid plans a choice of 403 or 402 Payment Required with a message, and Pay Per Crawl in limited rollout. A managed robots.txt can also publish your preferences without enforcing them.

One limit to know on Free: detection is by user agent string, so it catches the crawlers that announce themselves, which is all the major AI operators, and nothing that lies about its identity. Paid plans add Cloudflare's bot-management fingerprinting, which catches disguised scrapers. For finding out which assistants read your store, the Free detection is enough; for defending against bad actors, it is not.

04/12 Section

Reading The Crawlers Tab

Open the Crawlers tab and sort by requests. Then read it with the three categories in mind.

# For each row: crawler | operator | category | requests (period) | trend | allowed? 1. tag the category training / search / agent, from Cloudflare's label or the operator's crawler page 2. rank retrieval bots these are the ones that can cite you; note their request counts for section 8 3. find the surprises a training bot at the top of the list, a crawler you blocked still showing requests, an operator you have never heard of 4. check the trend a retrieval bot rising = you are being cited more; a training bot spiking = a model refresh and says nothing about you 5. note bandwidth Metrics tab; a crawler eating gigabytes on a small store is re-reading images, which section 9 covers # On Free: do this daily for a week to build a baseline, since the window is 24h.

The common surprise on ecommerce stores is a training crawler with more requests than any retrieval crawler, reading the full catalog including every variant URL. That is not a citation signal. It is a model collecting data, and whether you want it there is section 10's question. The second common surprise is a retrieval crawler that reads only the blog and never the product pages, which means the assistant is citing your guides and has no idea what you sell; section 9 covers that.

05/12 Section

Reading Robots.txt Compliance

The Robots.txt tab answers a question site owners used to have no way to ask: did the crawler I told to stay out of a path go there anyway. Cloudflare lists the crawler, its operator, the disallowed path it requested, the directive it violated and how many times.

Three things to do with that list.

  • Check your file first. If a crawler "violated" a directive, confirm the directive is written correctly; a typo in the user-agent token or a path pattern that does not match is the most common cause. The tab also shows whether your robots.txt returned 200, which is the second most common cause.
  • Separate the categories. Some operators run their training crawler and their user-triggered fetcher under different tokens with different rules; an agent fetch of a disallowed page on a user's behalf is a different thing from a training crawl of it. Read the operator's documentation before deciding it is bad faith.
  • Enforce where it matters. For a crawler that reads disallowed paths repeatedly, block it in the Crawlers tab, which enforces at the proxy rather than asking politely. Cloudflare is direct that robots.txt is a preference and AI Crawl Control is the wall.

Content Signals, the newer robots.txt directives that express preferences on training, search and AI input, show up here too. They are preferences, the same as any robots.txt line; the technical audit guide covers where to put them and what they do and do not accomplish.

06/12 Section

Without Cloudflare: Raw Server Logs

Every host keeps an access log. On shared WordPress hosting it is usually downloadable from the control panel (often under a logs or statistics section) or readable over SSH in the web server's log directory; on Shopify, the platform does not expose raw logs, so Cloudflare or an app is the path. Once you have the file, the analysis is a few lines of shell.

# 1. Count requests by AI crawler user agent (last 30 days of log) grep -iE 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|Google-Extended|Applebot-Extended|Bytespider|CCBot|Meta-ExternalAgent|Amazonbot' access.log \ | grep -oiE 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|Google-Extended|Applebot-Extended|Bytespider|CCBot|Meta-ExternalAgent|Amazonbot' \ | sort | uniq -c | sort -rn # 2. Which pages one crawler read most grep -i 'PerplexityBot' access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -50 # 3. Status codes that crawler received (403s = you are blocking it; 404s = it has stale URLs) grep -i 'ClaudeBot' access.log | awk '{print $9}' | sort | uniq -c | sort -rn # 4. Requests per day for one crawler (date field format varies; adjust the cut) grep -i 'GPTBot' access.log | cut -d[ -f2 | cut -d: -f1 | sort | uniq -c # 5. Did a blocked crawler request disallowed paths anyway grep -i 'Bytespider' access.log | grep -E '/cart|/checkout|/account' | wc -l

The field positions ($7 for path, $9 for status) match the common and combined log formats most hosts use; if yours differ, print one line and count. Keep the user agent list current from the crawler list guide, because new tokens appear and a grep that does not know a bot does not count it.

The raw log has one advantage over any dashboard: it is yours, it goes back as far as your host retains it, and it is not summarized. Pull it monthly into a folder and the history builds itself.

07/12 Section

Verifying A Crawler Is Who It Says It Is

A user agent string is a claim. Anyone can send "GPTBot" in a header, and scrapers do, because sites tend to allow the real one. Before you make a decision based on a crawler's behavior, confirm the traffic came from the operator.

# Most major operators publish the IP ranges their crawlers use; check the operator's crawler documentation page for the current list or JSON. 1. pull the IPs grep -i 'GPTBot' access.log | awk '{print $1}' | sort | uniq -c | sort -rn | head 2. compare against the operator's published ranges; a "GPTBot" from a residential ISP or a random VPS is not GPTBot 3. reverse DNS some operators support it: host <ip> should resolve to the operator's domain, and forward-resolving that name should return the same ip 4. on Cloudflare AI Crawl Control's crawler list is operator-verified for the known bots, so the fake ones land in general bot traffic rather than this tab # A spoofed crawler is a scraper. Treat it as one: rate limit or block by IP rather than by user agent.

This matters most when a crawler appears to ignore robots.txt. Half the time the "violation" is a spoofed user agent from a scraper that was never going to read robots.txt, and blocking the real operator over it costs you citations for nothing.

08/12 Section

The Crawl-To-Referral Ratio

For retrieval crawlers, one number decides whether the crawling is worth it: how many crawl requests it takes to produce one referred visitor. Crawls come from this post's method; referrals come from GA4 with a source-match channel group, which the GA4 tracking guide earlier in this series builds.

# Illustrative numbers for a mid-size content site; replace with your own from AI Crawl Control (or log counts) and a GA4 source-match Explore. These are not our figures. OPERATOR crawl requests / mo GA4 AI referral sessions / mo crawls per referral trend vs last month OpenAI 2,400 30 80 improving (was 110) Anthropic 900 15 60 improving (was 75) Perplexity 1,100 5 220 flat Google (AI) 3,000 not separable in GA4 (arrives as organic) n/a n/a Microsoft 600 2 300 worsening (was 200) # Reading it: a falling ratio = the assistant is citing you more per page read. A ratio in the thousands with no trend = it reads you and never sends anyone; look at WHICH pages (section 9) before blocking. # Training bots have no ratio. Do not put them in this table; decide them in section 10.

We do not print our own ratios here because the counts move monthly and a stale number would be quoted as a benchmark. The method is the point: the same operator, the same month, crawls over referrals. On a store with a healthy content cluster, retrieval crawlers cluster on the guides and the ratio improves as the cluster grows; on a store that is all product pages and no answers, the ratio stays high because there is nothing to cite. The citation audit covers the fixes on the content side.

09/12 Section

Which Pages They Read, And What That Tells You

The path list per crawler (command 2 in section 6, or the Metrics detail on paid Cloudflare plans) is the most actionable output of the whole exercise. Five patterns, and what each one means.

Pattern In The Path ListWhat It MeansWhat To Do
Retrieval bot reads guides, never product pagesThe assistant cites your content and does not know your catalogLink products from the guides; add Product schema; put a machine-readable catalog summary in llms.txt
Crawler hits thousands of variant and filter URLsFaceted navigation is generating crawl traps; it is wasting its budget and your bandwidthDisallow filter parameters in robots.txt; canonicalize variants; Cloudflare redirect rules for training bots
Many 404sThe crawler has stale URLs from an old crawl, a migration or a redesign301 the old paths; the crawler updates on its next pass
Many 403s to a bot you did not blockA CDN or WAF rule is challenging it; the assistant sees a blank pageFind the rule (Cloudflare Security Events shows which) and exempt verified crawlers
Same pages re-read dailyThose pages are being cited for live queries and the assistant checks them for freshnessKeep them current; they are your citation assets, and the landing-page report in GA4 should show them earning entries

The last row is the one worth building a routine around. The pages a retrieval crawler returns to are the pages assistants trust, and they map closely to the AI landing pages in GA4. Update them, add the next step and the offer, and link them to the guides you want cited next. The llms.txt guide covers the file that gives crawlers the catalog summary the first row is missing.

10/12 Section

Decide: Allow, Block, Or Charge

With the categories tagged, the ratios computed and the paths read, the decision is three rules.

  1. Retrieval and agent bots: allow, and fix what they cannot read. These are where citations and referrals come from. Blocking them makes you invisible to that assistant; challenging them makes you a blank page. If one has a terrible ratio, the fix is usually on your side (section 9), and blocking it is the last resort after the path list says it is reading nothing useful.
  2. Training bots: decide on principle, then enforce. They will never send a visitor, so analytics cannot decide this. If you want future models to know your brand and products, allow them and accept the bandwidth. If you object to your content training models, block them in AI Crawl Control (which enforces) and in robots.txt (which states the preference), and on a paid plan consider a 402 response with a licensing contact rather than a 403, which at least opens a conversation. Pay Per Crawl, where available, is the third option and still in limited rollout.
  3. Spoofed and unknown bots: treat as scrapers. Verify first (section 7). Anything claiming a known identity from the wrong IPs, or an unknown agent reading your whole catalog with no operator page, gets rate-limited or blocked by IP.

Whatever you decide, make robots.txt and the Cloudflare rules say the same thing. A crawler allowed in one and blocked in the other gets a conflicting signal, and the proxy wins, so the robots.txt preference is misleading everyone who reads it.

Free Resource

The Ecom Profit Box

Our library of ecommerce growth guides, including the AI search and technical audit frameworks behind this post.

Get It Free
30 Minutes

Get Your Crawler Report Read

Export your AI Crawl Control Crawlers tab or send a month of access logs. We will tag the categories, compute the ratios, find the blocked and stale paths, and hand you the allow-block list.

Book A Call
11/12 Section

The September 2026 Default Change

Cloudflare changed the default AI bot policies for newly onboarded domains effective September 15, 2026, and gave existing customers the option to opt out of the new defaults from Security Settings before that date. Domains already on Cloudflare keep their existing settings; new domains get the new defaults automatically. The July 2026 release that preceded it split AI bot controls into Search, Agent and Training categories with separate policies for each.

Why it belongs in a log-analysis post: if you moved a store onto Cloudflare recently, or launched a new domain, your default policy may be blocking a category you assumed was allowed, and the first place that shows up is a retrieval crawler receiving 403s in your logs. Check Security Settings for the AI bot policies on every domain, confirm they match the three rules in section 10, and re-check after any Cloudflare change to the defaults, which their changelog announces.

The general lesson holds for every host and CDN: bot protection that was configured for scrapers in 2024 is now also configured for the assistants you want reading you. The logs are how you find out.

12/12 Section

The Monthly Crawler Review

1. crawlers tab (or log count) every AI crawler, tagged training / search / agent, requests and trend 2. robots.txt tab file health, Content Signals present, violations by crawler and path 3. verify anomalies any violator or unknown agent checked against operator IP ranges 4. ratio table crawls / GA4 AI referrals per retrieval operator, with trend 5. path review top 50 paths per retrieval bot; the five patterns from section 9; the "re-read daily" list 6. status codes 403s to allowed bots (a rule is wrong), 404s (stale URLs to redirect) 7. policy check Cloudflare AI bot policies and robots.txt agree; new Cloudflare defaults reviewed 8. one action the single change with the biggest effect on the ratio: usually a page fix, sometimes a rule

The output is a short list of pages to update and, occasionally, one rule to change. It is not a big program, and the counterweight to selling it as one: for most stores, this is a half-hour a month for a founder or a marketer with dashboard access, and the value is in doing it every month, because the crawlers change faster than any guide about them.

GA4 tells you who came. The log tells you who read. Between the two is the ratio that says whether being read is worth it.
The whole post in three sentences
Key Takeaways

What To Remember

  • AI crawlers never appear in GA4 because they do not run JavaScript; the request log is the only record of which assistants read your store.
  • Sort every crawler into training, search/retrieval, or agent; only the second and third can ever send a visitor, and Cloudflare's July 2026 controls use the same three categories.
  • Cloudflare AI Crawl Control is on every plan, with Crawlers, Metrics and Robots.txt tabs; the Free plan shows 24 hours of metrics and detects by user agent only.
  • Without Cloudflare, a few grep and awk lines on the access log give the same counts, paths and status codes, with as much history as your host keeps.
  • Verify a crawler against its operator's published IP ranges before acting on its behavior; a spoofed user agent is a scraper, and half of apparent robots.txt violations are spoofs.
  • The crawl-to-referral ratio per retrieval operator, crawls over GA4 AI referrals, is the number that decides allow or block; training bots have no ratio and are decided on principle.
  • The pages a retrieval bot re-reads daily are your citation assets; keep them current, link products from them, and make robots.txt and Cloudflare rules agree.
Sources

Where This Came From

  1. Cloudflare, AI Crawl Control documentation, for the product's tabs, per-crawler controls, robots.txt tracking and Pay Per Crawl status, and the AI Crawl Control changelog for the Robots.txt tab, custom 402 and 403 block responses on paid plans, and always-free paths. Fetched September 15, 2026.
  2. o-c.do, How Do You Block AI Crawlers on Cloudflare, September 9, 2026, summarizing Cloudflare's statements on robots.txt being voluntary, Free-plan user-agent-only detection, the 24-hour metrics window, and the September 15, 2026 default change for new domains. An independent guide; Cloudflare's own Security Settings and changelog are authoritative.
  3. Operator crawler documentation pages (OpenAI, Anthropic, Perplexity, Google, Microsoft, Apple, Meta, ByteDance, Common Crawl) for user agent tokens, categories and published IP ranges. Cited by name; tokens and ranges change, and the crawler list guide on this site is maintained against them.
  4. Evolve Media Agency, Cloudflare and server-log audits on evolveamz.com and client stores, 2025-2026. First-party; the category framing, the ratio method, the five path patterns and the monthly review are ours. No first-party crawler counts are published in this post because they change monthly and would be quoted as benchmarks.

Questions

Twelve things store owners ask once they look at the log for the first time.
How do I see which AI crawlers are visiting my website?

Two ways. If your domain is on Cloudflare, open AI Crawl Control and read the Crawlers tab, which lists every known AI crawler that requested your site with operator, category, request count and trend; it is available on every plan. Otherwise download your web server's access log and count requests by user agent with grep (GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot and the rest), then list the paths each one requested.

Why do AI crawlers not show up in Google Analytics?

GA4 records visits through a JavaScript tag that runs in a browser. Crawlers request the raw HTML and do not execute the tag, so they leave no GA4 record at all. The only evidence of a crawler is the request log kept by your server or CDN. GA4 does show the referral visits that retrieval crawlers eventually produce, which is the other half of the ratio in this post.

What is the difference between GPTBot and OAI-SearchBot?

Different purposes from the same operator. GPTBot is a training crawler that collects pages for model training and produces no referral. OAI-SearchBot is a retrieval crawler that fetches current pages so ChatGPT can answer a live question and cite the source, which is where ChatGPT referrals come from. ChatGPT-User is a third token for fetches a user triggered directly. Each can be allowed or blocked separately.

What does Cloudflare AI Crawl Control do?

It shows which AI crawlers access your domain and how often (Crawlers tab), volume and bandwidth over time (Metrics tab), and which crawlers requested paths your robots.txt disallows (Robots.txt tab). It lets you allow or block each crawler at the proxy, which enforces rather than asks, and on paid plans return a 402 Payment Required with a licensing message. Formerly called AI Audit; on every plan as of September 2026.

Is the free Cloudflare plan enough for crawler analysis?

For finding out which assistants read your store, yes. Free detects crawlers by user agent, which catches every major AI operator because they identify themselves, and gives per-crawler allow and block. Its limits are a 24-hour metrics window (screenshot daily for history) and no fingerprinting of disguised bots, which paid plans add through bot management.

How do I know an AI crawler is real and not a spoofed user agent?

Pull the IP addresses behind the user agent from your log and compare them to the IP ranges the operator publishes on its crawler documentation page; several also support reverse DNS verification. A "GPTBot" arriving from a residential ISP or a random VPS is a scraper wearing a name. Cloudflare's crawler list is operator-verified for known bots, so spoofs land in general bot traffic instead.

Should I block AI training crawlers on my store?

It is a principle decision, because training crawlers never send a visitor and analytics cannot answer it. Allow them if you want future models to know your brand and products; block them if you object to your content training models. Whichever you choose, enforce it in Cloudflare (or your WAF) and state it in robots.txt so the two agree. Retrieval and agent crawlers are a different question and should generally stay allowed.

What is a crawl-to-referral ratio?

For one retrieval operator in one month, crawl requests from your log divided by AI referral sessions from GA4 for that operator's assistant. A falling ratio means the assistant is citing you more per page it reads. A very high ratio with no trend means it reads you and sends nobody, and the path list usually explains why. Training crawlers have no ratio because they never refer anyone.

What does it mean when a crawler reads my blog but not my product pages?

The assistant is citing your content and does not know what you sell. Fix it on your side: link products from the guides being read, add Product schema to product pages, and publish a machine-readable catalog summary in llms.txt so the crawler finds the products from the pages it already trusts.

Why is an allowed AI crawler getting 403 errors on my site?

A CDN, WAF or bot-protection rule is challenging or blocking it, and the assistant sees a blank page. On Cloudflare, Security Events shows which rule fired; exempt verified AI crawlers from it. Also check the AI bot policies in Security Settings, since Cloudflare changed defaults for newly onboarded domains on September 15, 2026, and a new domain may be blocking a category you assumed was open.

How do I stop AI crawlers from wasting bandwidth on filter and variant URLs?

Disallow the filter parameters in robots.txt, canonicalize variant URLs to the parent product, and on Cloudflare use redirect rules to steer training crawlers toward canonical content. A crawler hitting thousands of faceted URLs is a crawl trap; it is burning its budget on duplicates and your bandwidth on nothing, and the path list per crawler is how you spot it.

How often should I review AI crawler logs?

Monthly, for about thirty minutes: crawlers and trends, robots.txt violations verified against operator IPs, the ratio per retrieval operator, the top paths and status codes per crawler, and a check that Cloudflare policies and robots.txt agree. The output is a short list of pages to update and occasionally one rule to change. The value is in doing it every month, because the crawlers change faster than any guide about them.

Ian Smith, founder of Evolve Media Agency
Ian Smith
Founder, Evolve Media Agency

Ian founded Evolve Media Agency in 2017 and has spent a decade building Amazon, TikTok Shop and Shopify brands, including his own. The agency produces product photography, video, listing content, email and AI-search visibility work for ecommerce brands in the $1M to $10M range.

Read Ian's Story

Who Is Reading Your Store?

Export your Crawlers tab or send a month of access logs. In 30 minutes we will tag every crawler, compute the ratios, find the 403s and stale paths, and hand you the allow-block list and the pages to update.

3
Bot Categories, Three Rules