DECISION FRAMEWORK PUBLISHED AUGUST 19, 2026 · 14 MIN READ

Should You Block AI Crawlers?

On September 15, Cloudflare starts blocking mixed-use crawlers from ad-supported pages by default — and millions of site owners who never touched a bot setting are about to have this decision made for them. Six questions that determine whether an ecommerce brand should block, allow, or charge.

Sep 15Cloudflare default change lands
50.6%Of AI bot traffic is training crawlers
10.7%Is search bots, the ones that send clicks
6Questions in the framework
Quick Answer

For most ecommerce brands, no — you should not block AI crawlers, because your content is marketing for a product you sell rather than the product itself, and being absent from AI answers costs you more than any training-data contribution is worth. The calculation inverts for publishers whose content is the product, who have licensing leverage, and who monetize pageviews. The decision has become urgent rather than theoretical: on September 15, 2026, Cloudflare's default settings begin blocking mixed-use crawlers — bots that combine search indexing with AI training or agentic use — from any page carrying advertising. That default applies automatically to new customers, new sites, and all existing free-tier accounts, which means a very large number of site owners will have this decision made on their behalf unless they go and change a setting. Before September 15, check what your CDN is configured to do, then work the six questions in section six to decide deliberately rather than by default.

For two years this was a philosophical question you could postpone. On September 15 it becomes a setting that has a default, and defaults decide outcomes for the large majority of people who never open the panel.

The block-or-allow debate has been running since GPTBot got a name. Most of it has been abstract — arguments about fairness, about whether training on public content is transformative, about what publishers are owed. Genuinely interesting, and genuinely postponable if you had a business to run.

That changed on July 1, 2026, when Cloudflare announced that from September 15 its default configuration would start blocking mixed-use crawlers from ad-supported pages. The change applies automatically to new customers, to new sites created by existing customers, and to every existing free-tier account. No action required, no notification you are likely to read.

So the question is no longer whether you want to have an opinion about AI crawlers. It is whether your opinion or your CDN vendor's default is going to govern your site five weeks from now.

Definition

Mixed-use crawler — a bot that combines traditional search indexing with AI training or agentic retrieval under a single user-agent, making it impossible for a publisher to allow one function while refusing the other. Googlebot is the canonical example: it powers Google Search and feeds AI Overviews, with no separate token to distinguish them.

01/12SECTION

Why This Stopped Being Theoretical

Three things converged this year to move this from a debate to a deadline.

The traffic mix inverted

By June 2026, training crawlers accounted for roughly 50.6% of AI bot traffic on Cloudflare's network, while search bots — historically the ones that paid for access by sending clicks back — had fallen to about 10.7%. The bots consuming the most bandwidth are increasingly the ones returning the least.

The waste became visible

Cloudflare's data indicates more than half of AI crawl traffic goes to re-fetching pages that have not changed. That is a straightforward cost imposed on site owners with no corresponding benefit, and it is the kind of number that turns a philosophical position into an infrastructure argument.

Defaults moved

Opt-out protection at CDN scale changes the landscape faster than any individual decision could. Millions of ad-supported sites are about to be protected whether or not their owners ever thought about it — and a meaningful number of brands are about to lose citation surfaces they were relying on without understanding why.

Check this before September 15

If you are on Cloudflare, open the bot management settings and look at what your AI crawler controls are actually set to. The new default applies automatically to free-tier accounts and to any new site. Whatever you find, it should be a decision you made rather than one you inherited.

02/12SECTION

What the Cloudflare Default Actually Does

Precision matters here, because the headline coverage was loose about the scope.

ElementDetail
Effective dateSeptember 15, 2026
What is blockedMixed-use crawlers — bots blending search indexing with AI training or agent use
WhereOnly pages carrying advertising
Who it applies toNew Cloudflare customers, new sites from existing customers, and all existing free-tier accounts
Who it does not touchExisting paid customers who have already configured their own bot settings
Can you override itYes, manually, in your bot control settings
Stated intentPressure AI operators into separating search crawlers from training and agent crawlers

Cloudflare has also restructured its controls to divide crawlers into three categories — Search, Agent, and Training — so publishers can block or charge each independently. That taxonomy mirrors the one that should already be governing your robots.txt: the job a bot does, not the company that runs it, is what the rule should follow.

The Googlebot problem, again

The awkward part is that Googlebot is itself a mixed-use crawler. It powers Google Search and feeds AI Overviews through the same token, which is precisely why Cloudflare is pressuring operators to split their bots. Any default that catches mixed-use crawlers catches Googlebot, and for an ecommerce brand losing Googlebot is not a trade-off worth entertaining. Apple, Google and Microsoft each offer AI opt-out mechanisms that may let their crawlers stay in scope, but if you carry ads and you are on a default configuration, this is worth verifying rather than assuming.

03/12SECTION

The Crawl-to-Referral Ratios Behind It

The numbers driving publisher anger are worth understanding even if you land on the allow side, because they are the strongest version of the opposing argument.

Reporting on the Cloudflare announcement put Anthropic's crawler at roughly 38,000 page fetches for every one referral visit sent back to a publisher. OpenAI's ratio came in nearer 1,091 crawls per referral. Both are enormous compared with traditional search, where the implicit bargain was that indexing bought you traffic.

Read the metric carefully

Those figures measure crawls per referral visit, not per appearance in an answer. Being named inside an AI response has real brand value that produces no click and therefore appears nowhere in that ratio. For a brand, an uncited-but-recommended mention can still generate a sale later. For an ad-supported publisher, it monetizes at exactly zero.

That distinction is the entire fight, and it is also why the correct answer differs so sharply between publishers and ecommerce brands. If your revenue comes from impressions on your own pages, a visit is the only thing that pays. If your revenue comes from selling a product, being recommended is the thing that pays and the visit is incidental.

An appearance in an AI answer does something for a brand and nothing for an ad stack. That single asymmetry explains why publishers and ecommerce brands should reach opposite conclusions from identical data.
Ian Smith · Evolve Media Agency
04/12SECTION

The Case For Blocking, Steelmanned

Presented as its strongest advocates would put it, without hedging.

  • You are subsidizing a competitor. Every page an AI system ingests improves a product that increasingly answers the questions your content was written to answer, without sending anyone to you. You are paying bandwidth to reduce your own future traffic.
  • The exchange is not reciprocal. Search crawling came with an implicit bargain: index my content, send me visitors. At thousands of crawls per referral, that bargain has collapsed and continuing to honor it is habit rather than strategy.
  • Blocking creates negotiating leverage. Content that cannot be taken freely can be licensed. Several large publishers have converted blocks into paid agreements, and no one negotiates from a position of already having given the thing away.
  • The infrastructure cost is real. More than half of AI crawl traffic re-fetches unchanged pages. You are paying for compute and bandwidth that produces nothing for anyone.
  • Consent should be affirmative. A default of open access to any commercial actor who shows up with a user-agent string is a strange norm for any other asset a business owns.
  • Blocking is reversible. You can unblock tomorrow. Content already absorbed into a training run cannot be recalled.

That last point is the sharpest one, and it deserves more weight than it usually gets. The asymmetry between a reversible block and an irreversible ingestion is a genuine argument for caution rather than openness.

05/12SECTION

The Case For Allowing, Steelmanned

Same treatment for the other side.

  • Absence is not neutrality. If you block, the question still gets answered — using your competitor's content. You have not withheld yourself from the conversation, you have removed yourself from the shortlist.
  • Your content is marketing, not inventory. For an ecommerce brand, a blog post exists to sell a product. It has no standalone licensing value. Withholding it protects an asset that does not exist while forfeiting distribution that does.
  • AI referral traffic converts far better. The visitor arrives having been recommended rather than having found you in a list of ten. That quality premium is well documented and it applies to a channel you would be opting out of entirely.
  • Licensing leverage is concentrated, not distributed. Major publishers with distinctive archives can negotiate. A brand with 200 posts about product selection is not going to be offered a deal, so the leverage argument is borrowing a benefit that will not arrive.
  • Citations compound. Entity recognition and topical authority build slowly and non-linearly. Blocking during the establishment phase means competitors become the corroborated default, and displacing an established entity is far harder than establishing one.
  • The retrieval layer is separable. You do not have to choose between all-in and all-out. Training and retrieval run on distinct user-agents, so the actual choice space is wider than the debate implies.
06/12SECTION

The Six-Question Decision Framework

Work these in order. The first three usually settle it.

THE SIX QUESTIONSWORK IN ORDER
QUESTION 01
Is Your Content The Product?

If people pay for access to your content, blocking is defensible. If your content exists to sell something else, blocking protects an asset that does not exist.

QUESTION 02
How Do You Monetize A Visit?

Ad impressions require a pageview and die without it. Product sales survive an uncited recommendation. This single answer flips the calculation.

QUESTION 03
Do You Have Real Leverage?

Would an AI company notice your absence? Distinctive archives and proprietary data have leverage. A well-executed brand blog does not, and pretending otherwise costs you distribution for nothing.

QUESTION 04
What Is Your Current AI Share?

Measure before deciding. If AI referrals are already meaningful, blocking has a known price. If they are zero, find out whether that is because you are blocked before concluding the channel is worthless.

QUESTION 05
What Is Your Competitive Set Doing?

If competitors are open and you block, you hand them the category. If the whole category blocks, models answer generically and nobody wins. Check rather than assume.

QUESTION 06
What Is Your 24-Month Bet?

If you expect AI-mediated discovery to grow, entity establishment now is cheap and later is expensive. If you expect it to plateau, the urgency drops considerably.

How the answers usually resolve

If you answered no to question one and product sales to question two, you are an ecommerce brand and the answer is almost certainly allow retrieval, decide training on policy. If you answered yes to one and ad impressions to two, you are a publisher and blocking or charging is defensible. Question three is the honesty check: most businesses that believe they have licensing leverage do not.

07/12SECTION

Publishers vs Ecommerce Brands

Most of the writing on this question is by publishers, for publishers, and it is being read by ecommerce operators as though it applies to them. It largely does not.

DimensionPublisherEcommerce Brand
What content isThe productMarketing for the product
How a visit paysAd impression, requires the pageviewProduct sale, survives an uncited mention
Uncited recommendationWorth nothingWorth a great deal
Licensing leverageReal for large archivesEffectively none
Cost of invisibilityLost impressionsLost consideration-set membership
Reasonable defaultBlock or chargeAllow retrieval, decide training on policy

The confusion is understandable. Publishers write more, write faster, and have a direct interest in the outcome. But an ecommerce brand adopting a publisher's posture is copying a strategy built on an economics it does not share.

08/12SECTION

The Middle Path Most Brands Should Take

Because training and retrieval run on separate user-agents, the practical choice is not binary. The configuration most ecommerce brands should land on:

  • Allow every retrieval crawler unconditionally. These are how citations happen. Blocking them removes you from AI answers with no offsetting benefit.
  • Allow every user-triggered fetcher. Blocking these only breaks the moment a real person tries to hand your page to their assistant.
  • Decide training collectors on policy. Most brands leave them open and lose nothing. Blocking them is defensible and costs little. This is the one genuine judgment call.
  • Block the non-compliant. Bytespider and similar high-volume, low-return crawlers, enforced at the server rather than in robots.txt.
  • Protect the paths that should never be crawled. Cart, checkout, account, admin. This has always been true and has nothing to do with AI.
  • Verify your CDN agrees. A permissive robots.txt means nothing if Cloudflare is challenging crawlers above it. After September 15 this check moves from good practice to necessary.

The implementation detail for each of those layers is covered in our AI crawler list for ecommerce, and the diagnostic for finding what is silently blocking you is in the AI crawler technical audit.

FREE 30-MINUTE CALL

Check your config before September 15

We will look at your robots.txt, your Cloudflare bot settings, and your current AI referral share together, and tell you what the default change will do to you.

Book a Strategy Call →
FREE RESOURCE

The Ecom Profit Box

Eleven playbooks on listings, conversion, images, and email. Built for operators, no fluff, no email sequence.

Grab It Free →
09/12SECTION

Pay Per Use and Whether To Engage

Cloudflare has evolved its Pay Per Crawl marketplace into Pay Per Use, which compensates publishers when their content creates value inside an AI answer rather than merely when a bot fetches the page. That is a meaningful design change — it addresses the central complaint about flat per-crawl pricing, which was that fetches and value are only loosely related.

Cloudflare has also announced partnerships intended to compensate publishers when content appears in search results or agent queries, and a business insights dashboard surfacing referral volumes and crawl-to-referral ratios per crawler — data specifically useful for licensing negotiations.

Should an ecommerce brand engage with it?

Probably not yet, and the reasoning is unglamorous. Monetization frameworks reward content with standalone value: original research, distinctive archives, proprietary data. A brand blog optimized to sell products has little of that. The realistic outcome of engaging is administrative overhead in exchange for a payment that rounds to nothing, while the citation value you were already getting for free is what actually moves your revenue.

Where it becomes genuinely worth watching is if you publish original data nobody else has — industry benchmarks, proprietary survey results, category research. That content has real standalone value and the calculation changes.

The part worth taking seriously

Ignore the payments and pay attention to the reporting. Per-crawler referral volumes and crawl-to-referral ratios are exactly the data most brands currently lack. Knowing which AI systems consume your content and which actually send anyone back is useful whether or not you ever charge for access.

10/12SECTION

What Actually Happens After You Block

A realistic timeline, because the consequences are slower and quieter than people expect — which is exactly what makes them easy to misattribute.

WindowWhat you observeWhat is happening
Days 1–7NothingIndexes still hold your content; answers unchanged
Weeks 2–6Still mostly nothingCached content ages out gradually rather than dropping
Months 2–4Citations thin outFresh content never enters the index; older mentions persist
Months 4–8Competitors named insteadYour slot in the answer gets filled by whoever stayed open
Months 8+Absent from the categoryEntity association decays; you are no longer a default answer
After unblockingSlow returnRe-establishing an entity takes materially longer than losing it

The dangerous feature of this curve is the flat start. Nothing visible happens for six weeks, which reads as confirmation the block was costless. The cost arrives in month four, by which point almost nobody connects it to a robots.txt change made a quarter earlier. Our analysis of how AI citations compound over twelve months covers the same curve running in the other direction.

11/12SECTION

The Risk Nobody Is Pricing: Other People's Blocks

Here is the second-order effect that gets almost no attention, and it affects you even if you never touch a setting.

Your citations do not only come from your own pages. They come from roundups that mention you, forum threads that recommend you, review sites that cover you, and industry publications that quote you. When those sites block AI crawlers — or get blocked by a CDN default they never configured — the pages that were carrying your brand into AI answers disappear from the retrievable corpus.

You did nothing. Your configuration is unchanged. Your visibility drops anyway.

What this implies practically

  • Owned content matters more than it used to. Pages you control cannot be blocked by someone else's default.
  • Diversify where your mentions live. If your citation profile depends heavily on two or three ad-supported publishers, that is now a concentration risk with a date on it.
  • Re-baseline after mid-September. If your AI visibility moves in October, third-party blocks are a candidate explanation before anything you did.
  • Platforms with different economics are more durable. Community sites, forums, and non-ad-supported sources are less exposed to this particular change.

This is the strongest available argument for building a deep owned-content base rather than relying on earned mentions, and it is covered further in our piece on building a compounding content moat.

12/12SECTION

Decision Matrix by Business Model

The summary answer for each situation, with the reasoning compressed.

Business modelTraining botsRetrieval botsWhy
Ecommerce brandAllow, or block on policyAllowContent is marketing; citations sell product
DTC with content engineAllowAllowEntity establishment is the whole play
Local service businessAllowAllowThin corpus means every source counts
SaaS / B2BPolicy choiceAllowDocs and comparisons drive high-intent discovery
Ad-supported publisherBlock or chargeEvaluate carefullyPageview is the unit of revenue
Subscription publisherBlockAllow teasers onlyContent is the product being sold
Original research / dataBlockAllow with attributionGenuine licensing leverage exists
Marketplace sellerAllowAllowProduct discovery happens off your domain anyway

The three things to do before September 15

  1. Check your CDN bot settings. Find out what is configured today. If you are on a free tier or a recently created site, the default is about to change under you.
  2. Measure your current AI referral share. You cannot evaluate the cost of a block without knowing what you are giving up. The AI visibility audit guide covers the diagnostic, and several tracking tools automate the sampling.
  3. Write the decision down. Whatever you choose, record why, so that in four months when something moves you can tell whether this was the cause.

The worst outcome is not blocking or allowing. It is discovering in December that a setting you never saw made the decision for you in September.

Key Takeaways

The Short Version

  • From September 15, 2026, Cloudflare's defaults block mixed-use crawlers from ad-supported pages, applying automatically to new customers, new sites, and all existing free-tier accounts.
  • By June 2026 training crawlers were roughly 50.6% of AI bot traffic on Cloudflare's network while search bots had fallen to about 10.7%, and over half of AI crawl traffic re-fetches unchanged pages.
  • Crawl-to-referral ratios reported around 38,000 to 1 for Anthropic and 1,091 to 1 for OpenAI measure referral visits, not answer appearances — which is why publishers and brands should reach opposite conclusions.
  • For most ecommerce brands the answer is allow retrieval unconditionally and decide training on policy, because content is marketing rather than inventory and you have no real licensing leverage.
  • Blocking has a flat consequence curve: nothing visible for six weeks, citations thinning at months two to four, competitors occupying your slot by months four to eight.
  • The unpriced risk is other people's blocks. When sites that mention you go dark to crawlers, your visibility falls without you changing anything.
  • Before September 15: check your CDN bot settings, measure your current AI referral share, and write down the decision and the reasoning.

Common Questions

Blocking AI Crawlers
FAQ

Should an ecommerce brand block AI crawlers?

Generally no. Your content is marketing for a product you sell rather than the product itself, so it has no standalone licensing value to protect, and being absent from AI answers costs you consideration-set membership at the exact moment buyers are deciding. The defensible middle path is to allow every retrieval and user-triggered crawler unconditionally and treat training collectors as a separate policy decision, since those two layers run on different user-agents.

What is Cloudflare changing on September 15, 2026?

Its default settings begin blocking mixed-use crawlers — bots that combine search indexing with AI training or agentic use — from any page carrying advertising. The change applies automatically to new Cloudflare customers, new sites created by existing customers, and all existing free-tier accounts. Site owners can override it manually. The stated intent is to pressure AI operators into separating their search crawlers from their training and agent crawlers.

Does the Cloudflare change affect my site if I do not run ads?

The default specifically targets pages that carry advertising, so a site with no ad units is outside the stated scope of that particular default. But this is worth verifying rather than assuming, because bot management settings interact in ways that are not always obvious, and other Cloudflare bot controls can challenge or block AI crawlers independently of this change. Open your bot settings and look at what is actually configured.

What is a mixed-use crawler?

A bot that serves both traditional search indexing and AI training or agentic retrieval under a single user-agent, which makes it impossible for a publisher to permit one function while refusing the other. Googlebot is the clearest example: it powers Google Search and simultaneously feeds AI Overviews with no separate token distinguishing them. Cloudflare's stated goal in blocking mixed-use crawlers by default is to push operators into splitting these functions into separate identifiable bots.

What are crawl-to-referral ratios and why do they matter?

They measure how many pages a crawler fetches for each visitor it sends back. Reporting around the Cloudflare announcement put Anthropic's crawler near 38,000 crawls per referral and OpenAI's around 1,091. The important caveat is that the metric counts referral visits, not appearances in answers. A brand mentioned in an AI response without a click still gains something; an ad-supported publisher gains nothing. That asymmetry is why the two should reach different conclusions.

If I block AI crawlers, how quickly will I notice?

You will not, for about six weeks, and that flat start is the trap. Existing indexes still hold your content and cached material ages out gradually. Citations thin around months two to four as fresh content never enters the index, competitors occupy your slot by months four to eight, and by month eight or later your entity association with the category has meaningfully decayed. Almost nobody connects a month-four decline to a configuration change made a quarter earlier.

Can I block training but still get cited in AI answers?

Yes, because training collection and retrieval indexing run on separate user-agents at the major labs. Blocking GPTBot and ClaudeBot while allowing OAI-SearchBot, Claude-SearchBot and PerplexityBot opts you out of training corpora while keeping you eligible for citation. The honest caveat is that this boundary is a policy commitment those organizations defined themselves, not a technical wall, so if your opt-out needs to be legally durable then robots.txt is not the right instrument.

Is Pay Per Use worth engaging with as an ecommerce brand?

Probably not for the payments. Monetization frameworks reward content with standalone value such as original research, distinctive archives, or proprietary data, and a brand blog written to sell products has little of that. The realistic outcome is administrative overhead for revenue that rounds to nothing. What is worth paying attention to is the reporting: per-crawler referral volumes and crawl-to-referral ratios are exactly the data most brands lack, regardless of whether you ever charge for access.

What if my competitors block and I do not?

That is a favorable position, not a risky one. If competitors remove themselves from the retrievable corpus and you stay open, you become the available source for category questions and your citation share rises without any additional work. The reverse is the danger. If your category is broadly open and you alone block, models simply name the businesses that stayed visible and you have handed over the category.

Can other websites blocking AI crawlers hurt my visibility?

Yes, and this is the most overlooked risk in the whole topic. Your citations come partly from roundups, forum threads, review sites and publications that mention you. When those sites block crawlers — or inherit a CDN default they never configured — the pages carrying your brand into AI answers leave the retrievable corpus. Your configuration is unchanged and your visibility falls anyway, which is a strong argument for building owned content that nobody else can block.

Is blocking reversible?

The setting is instantly reversible; the consequences are not symmetric. You can unblock tomorrow, but re-establishing entity association takes materially longer than losing it, because a competitor has usually occupied your position in the meantime and displacing an established entity is harder than establishing one. Conversely, content already absorbed into a completed training run cannot be recalled, which is the strongest argument on the blocking side.

What should I actually do before September 15?

Three things. Check your CDN bot settings and find out what is configured today, especially if you are on a free tier or a recently created site where the default is about to change. Measure your current AI referral share so you know what a block would actually cost. Then write the decision and the reasoning down, so that if something moves in four months you can tell whether this was the cause rather than guessing.

Ian Smith, Founder of Evolve Media Agency
Ian Smith
Founder, Evolve Media Agency · AI Search & Ecommerce Specialist

Ian co-founded Evolve Media Agency in 2017 with his wife Megan. Over 9 years he has worked with $1M-$10M ecommerce brands on AI search visibility, schema infrastructure, content production, and channel diversification. Based in Colorado. Read Ian’s full bio →

Work With Ian

Decide it, do not inherit it

Before September 15.

Book a strategy call and we will check your robots.txt, your CDN bot settings, and your current AI referral share, then tell you exactly what the default change does to your site.