An assistant that cites you had to read you first. The reading is in your logs, dated, by page, and most stores have never looked.
The AI crawler list on this site names the bots and what each one does. This is the follow-up it promised: how to see them on your own store. It uses Cloudflare because that is what we run and what most Shopify and WordPress stores in our client base run, and it covers raw logs for everyone else. Nothing here needs a developer; it needs a dashboard login and, for the raw-log path, a terminal and ten minutes.
Evolve Media Agency sells AI search visibility work, which includes exactly this kind of audit. Cloudflare's features and interface labels are described as documented on September 15, 2026, and Cloudflare changes them often; the docs linked in the sources are current where this post is not. We have no relationship with Cloudflare beyond being a customer.
Why This Is A Log Question, Not An Analytics One
GA4 fires from a JavaScript tag in the browser. Crawlers request the HTML and leave; they do not run the tag. So the assistant that read forty of your product pages last night left no trace in GA4, and the only record is the web server's or the CDN's request log: timestamp, IP, user agent, path, status code, bytes.
That record answers questions analytics never can. Which assistants know your catalog. How often they refresh it. Which pages they think are worth re-reading. Whether the crawler you blocked in robots.txt kept crawling. And, joined to GA4's referral data, whether the reading ever turns into a visitor, which is the number that decides whether to keep letting it read.
AI crawler. A bot operated by an AI company that requests pages from your site for one of three purposes: to collect training data for a model, to retrieve current content so an assistant can answer a question and cite the source, or to fetch a page an assistant's user asked it to open. Each identifies itself with a user agent string (GPTBot, ClaudeBot, PerplexityBot and so on), most publish the IP ranges they crawl from, and their behavior toward robots.txt varies by operator and purpose.
The Three Kinds Of AI Bot
Every decision later in this post depends on which category a crawler is in, and Cloudflare adopted the same three-way split in its July 2026 bot controls.
| Category | What It Does | Examples (User Agents) | Sends Referrals? | Default Stance For A Store |
|---|---|---|---|---|
| Training | Collects pages to train or fine-tune a model; nothing you see later is tied to a visit | GPTBot, ClaudeBot, Google-Extended (a control token rather than a bot), CCBot, Applebot-Extended, Meta-ExternalAgent, Bytespider | No | A principle decision: allow if you want future models to know your brand; block if you object to training use |
| Search / retrieval | Fetches current pages to answer a live question and cite the source | OAI-SearchBot, PerplexityBot, Claude-SearchBot, Bingbot (for Copilot), Googlebot (for AI Overviews and AI Mode) | Yes, this is where AI referral traffic comes from | Allow, and measure the ratio |
| Agent / user-triggered | Fetches a specific page because a user asked the assistant to open, summarize or shop it | ChatGPT-User, Claude-User, Perplexity-User, and browser-agent traffic | Sometimes; often the user follows through | Allow; blocking this blocks a customer |
The operator names shift and new bots appear monthly; the categories do not. When a user agent you do not recognize shows up, the first question is which of the three it is, and the operator's published crawler page usually says. The robots.txt guide has the current token for each and the directive that controls it.
Cloudflare AI Crawl Control: What Each Tab Shows
AI Crawl Control (the product formerly called AI Audit) is in the Cloudflare dashboard under your domain and, as of September 2026, is available on every plan including Free. It has three views that matter for analysis and one for action.
Every known AI crawler that requested your domain, with operator, category, request count and trend, and a per-crawler allow or block control. This is the list you are here for.
Volume, bandwidth and, on paid plans, request-to-referral patterns over time. On Free the window is the last 24 hours; paid plans get longer ranges. Screenshot it daily if you are on Free and want history.
Health of your robots.txt (status code, whether one exists per hostname), whether it carries Content Signals, and, the useful part, which crawlers requested paths your robots.txt disallows, with the path and the count.
Per-crawler block and allow on every plan; on paid plans a choice of 403 or 402 Payment Required with a message, and Pay Per Crawl in limited rollout. A managed robots.txt can also publish your preferences without enforcing them.
One limit to know on Free: detection is by user agent string, so it catches the crawlers that announce themselves, which is all the major AI operators, and nothing that lies about its identity. Paid plans add Cloudflare's bot-management fingerprinting, which catches disguised scrapers. For finding out which assistants read your store, the Free detection is enough; for defending against bad actors, it is not.
Reading The Crawlers Tab
Open the Crawlers tab and sort by requests. Then read it with the three categories in mind.
The common surprise on ecommerce stores is a training crawler with more requests than any retrieval crawler, reading the full catalog including every variant URL. That is not a citation signal. It is a model collecting data, and whether you want it there is section 10's question. The second common surprise is a retrieval crawler that reads only the blog and never the product pages, which means the assistant is citing your guides and has no idea what you sell; section 9 covers that.
Reading Robots.txt Compliance
The Robots.txt tab answers a question site owners used to have no way to ask: did the crawler I told to stay out of a path go there anyway. Cloudflare lists the crawler, its operator, the disallowed path it requested, the directive it violated and how many times.
Three things to do with that list.
- Check your file first. If a crawler "violated" a directive, confirm the directive is written correctly; a typo in the user-agent token or a path pattern that does not match is the most common cause. The tab also shows whether your robots.txt returned 200, which is the second most common cause.
- Separate the categories. Some operators run their training crawler and their user-triggered fetcher under different tokens with different rules; an agent fetch of a disallowed page on a user's behalf is a different thing from a training crawl of it. Read the operator's documentation before deciding it is bad faith.
- Enforce where it matters. For a crawler that reads disallowed paths repeatedly, block it in the Crawlers tab, which enforces at the proxy rather than asking politely. Cloudflare is direct that robots.txt is a preference and AI Crawl Control is the wall.
Content Signals, the newer robots.txt directives that express preferences on training, search and AI input, show up here too. They are preferences, the same as any robots.txt line; the technical audit guide covers where to put them and what they do and do not accomplish.
Without Cloudflare: Raw Server Logs
Every host keeps an access log. On shared WordPress hosting it is usually downloadable from the control panel (often under a logs or statistics section) or readable over SSH in the web server's log directory; on Shopify, the platform does not expose raw logs, so Cloudflare or an app is the path. Once you have the file, the analysis is a few lines of shell.
The field positions ($7 for path, $9 for status) match the common and combined log formats most hosts use; if yours differ, print one line and count. Keep the user agent list current from the crawler list guide, because new tokens appear and a grep that does not know a bot does not count it.
The raw log has one advantage over any dashboard: it is yours, it goes back as far as your host retains it, and it is not summarized. Pull it monthly into a folder and the history builds itself.
Verifying A Crawler Is Who It Says It Is
A user agent string is a claim. Anyone can send "GPTBot" in a header, and scrapers do, because sites tend to allow the real one. Before you make a decision based on a crawler's behavior, confirm the traffic came from the operator.
This matters most when a crawler appears to ignore robots.txt. Half the time the "violation" is a spoofed user agent from a scraper that was never going to read robots.txt, and blocking the real operator over it costs you citations for nothing.
The Crawl-To-Referral Ratio
For retrieval crawlers, one number decides whether the crawling is worth it: how many crawl requests it takes to produce one referred visitor. Crawls come from this post's method; referrals come from GA4 with a source-match channel group, which the GA4 tracking guide earlier in this series builds.
We do not print our own ratios here because the counts move monthly and a stale number would be quoted as a benchmark. The method is the point: the same operator, the same month, crawls over referrals. On a store with a healthy content cluster, retrieval crawlers cluster on the guides and the ratio improves as the cluster grows; on a store that is all product pages and no answers, the ratio stays high because there is nothing to cite. The citation audit covers the fixes on the content side.
Which Pages They Read, And What That Tells You
The path list per crawler (command 2 in section 6, or the Metrics detail on paid Cloudflare plans) is the most actionable output of the whole exercise. Five patterns, and what each one means.
| Pattern In The Path List | What It Means | What To Do |
|---|---|---|
| Retrieval bot reads guides, never product pages | The assistant cites your content and does not know your catalog | Link products from the guides; add Product schema; put a machine-readable catalog summary in llms.txt |
| Crawler hits thousands of variant and filter URLs | Faceted navigation is generating crawl traps; it is wasting its budget and your bandwidth | Disallow filter parameters in robots.txt; canonicalize variants; Cloudflare redirect rules for training bots |
| Many 404s | The crawler has stale URLs from an old crawl, a migration or a redesign | 301 the old paths; the crawler updates on its next pass |
| Many 403s to a bot you did not block | A CDN or WAF rule is challenging it; the assistant sees a blank page | Find the rule (Cloudflare Security Events shows which) and exempt verified crawlers |
| Same pages re-read daily | Those pages are being cited for live queries and the assistant checks them for freshness | Keep them current; they are your citation assets, and the landing-page report in GA4 should show them earning entries |
The last row is the one worth building a routine around. The pages a retrieval crawler returns to are the pages assistants trust, and they map closely to the AI landing pages in GA4. Update them, add the next step and the offer, and link them to the guides you want cited next. The llms.txt guide covers the file that gives crawlers the catalog summary the first row is missing.
Decide: Allow, Block, Or Charge
With the categories tagged, the ratios computed and the paths read, the decision is three rules.
- Retrieval and agent bots: allow, and fix what they cannot read. These are where citations and referrals come from. Blocking them makes you invisible to that assistant; challenging them makes you a blank page. If one has a terrible ratio, the fix is usually on your side (section 9), and blocking it is the last resort after the path list says it is reading nothing useful.
- Training bots: decide on principle, then enforce. They will never send a visitor, so analytics cannot decide this. If you want future models to know your brand and products, allow them and accept the bandwidth. If you object to your content training models, block them in AI Crawl Control (which enforces) and in robots.txt (which states the preference), and on a paid plan consider a 402 response with a licensing contact rather than a 403, which at least opens a conversation. Pay Per Crawl, where available, is the third option and still in limited rollout.
- Spoofed and unknown bots: treat as scrapers. Verify first (section 7). Anything claiming a known identity from the wrong IPs, or an unknown agent reading your whole catalog with no operator page, gets rate-limited or blocked by IP.
Whatever you decide, make robots.txt and the Cloudflare rules say the same thing. A crawler allowed in one and blocked in the other gets a conflicting signal, and the proxy wins, so the robots.txt preference is misleading everyone who reads it.
The Ecom Profit Box
Our library of ecommerce growth guides, including the AI search and technical audit frameworks behind this post.
Get It FreeGet Your Crawler Report Read
Export your AI Crawl Control Crawlers tab or send a month of access logs. We will tag the categories, compute the ratios, find the blocked and stale paths, and hand you the allow-block list.
Book A CallThe September 2026 Default Change
Cloudflare changed the default AI bot policies for newly onboarded domains effective September 15, 2026, and gave existing customers the option to opt out of the new defaults from Security Settings before that date. Domains already on Cloudflare keep their existing settings; new domains get the new defaults automatically. The July 2026 release that preceded it split AI bot controls into Search, Agent and Training categories with separate policies for each.
Why it belongs in a log-analysis post: if you moved a store onto Cloudflare recently, or launched a new domain, your default policy may be blocking a category you assumed was allowed, and the first place that shows up is a retrieval crawler receiving 403s in your logs. Check Security Settings for the AI bot policies on every domain, confirm they match the three rules in section 10, and re-check after any Cloudflare change to the defaults, which their changelog announces.
The general lesson holds for every host and CDN: bot protection that was configured for scrapers in 2024 is now also configured for the assistants you want reading you. The logs are how you find out.
The Monthly Crawler Review
The output is a short list of pages to update and, occasionally, one rule to change. It is not a big program, and the counterweight to selling it as one: for most stores, this is a half-hour a month for a founder or a marketer with dashboard access, and the value is in doing it every month, because the crawlers change faster than any guide about them.
GA4 tells you who came. The log tells you who read. Between the two is the ratio that says whether being read is worth it.
What To Remember
- AI crawlers never appear in GA4 because they do not run JavaScript; the request log is the only record of which assistants read your store.
- Sort every crawler into training, search/retrieval, or agent; only the second and third can ever send a visitor, and Cloudflare's July 2026 controls use the same three categories.
- Cloudflare AI Crawl Control is on every plan, with Crawlers, Metrics and Robots.txt tabs; the Free plan shows 24 hours of metrics and detects by user agent only.
- Without Cloudflare, a few grep and awk lines on the access log give the same counts, paths and status codes, with as much history as your host keeps.
- Verify a crawler against its operator's published IP ranges before acting on its behavior; a spoofed user agent is a scraper, and half of apparent robots.txt violations are spoofs.
- The crawl-to-referral ratio per retrieval operator, crawls over GA4 AI referrals, is the number that decides allow or block; training bots have no ratio and are decided on principle.
- The pages a retrieval bot re-reads daily are your citation assets; keep them current, link products from them, and make robots.txt and Cloudflare rules agree.
Where This Came From
- Cloudflare, AI Crawl Control documentation, for the product's tabs, per-crawler controls, robots.txt tracking and Pay Per Crawl status, and the AI Crawl Control changelog for the Robots.txt tab, custom 402 and 403 block responses on paid plans, and always-free paths. Fetched September 15, 2026.
- o-c.do, How Do You Block AI Crawlers on Cloudflare, September 9, 2026, summarizing Cloudflare's statements on robots.txt being voluntary, Free-plan user-agent-only detection, the 24-hour metrics window, and the September 15, 2026 default change for new domains. An independent guide; Cloudflare's own Security Settings and changelog are authoritative.
- Operator crawler documentation pages (OpenAI, Anthropic, Perplexity, Google, Microsoft, Apple, Meta, ByteDance, Common Crawl) for user agent tokens, categories and published IP ranges. Cited by name; tokens and ranges change, and the crawler list guide on this site is maintained against them.
- Evolve Media Agency, Cloudflare and server-log audits on evolveamz.com and client stores, 2025-2026. First-party; the category framing, the ratio method, the five path patterns and the monthly review are ours. No first-party crawler counts are published in this post because they change monthly and would be quoted as benchmarks.

