For two years this was a philosophical question you could postpone. On September 15 it becomes a setting that has a default, and defaults decide outcomes for the large majority of people who never open the panel.
The block-or-allow debate has been running since GPTBot got a name. Most of it has been abstract — arguments about fairness, about whether training on public content is transformative, about what publishers are owed. Genuinely interesting, and genuinely postponable if you had a business to run.
That changed on July 1, 2026, when Cloudflare announced that from September 15 its default configuration would start blocking mixed-use crawlers from ad-supported pages. The change applies automatically to new customers, to new sites created by existing customers, and to every existing free-tier account. No action required, no notification you are likely to read.
So the question is no longer whether you want to have an opinion about AI crawlers. It is whether your opinion or your CDN vendor's default is going to govern your site five weeks from now.
Mixed-use crawler — a bot that combines traditional search indexing with AI training or agentic retrieval under a single user-agent, making it impossible for a publisher to allow one function while refusing the other. Googlebot is the canonical example: it powers Google Search and feeds AI Overviews, with no separate token to distinguish them.
Why This Stopped Being Theoretical
Three things converged this year to move this from a debate to a deadline.
The traffic mix inverted
By June 2026, training crawlers accounted for roughly 50.6% of AI bot traffic on Cloudflare's network, while search bots — historically the ones that paid for access by sending clicks back — had fallen to about 10.7%. The bots consuming the most bandwidth are increasingly the ones returning the least.
The waste became visible
Cloudflare's data indicates more than half of AI crawl traffic goes to re-fetching pages that have not changed. That is a straightforward cost imposed on site owners with no corresponding benefit, and it is the kind of number that turns a philosophical position into an infrastructure argument.
Defaults moved
Opt-out protection at CDN scale changes the landscape faster than any individual decision could. Millions of ad-supported sites are about to be protected whether or not their owners ever thought about it — and a meaningful number of brands are about to lose citation surfaces they were relying on without understanding why.
If you are on Cloudflare, open the bot management settings and look at what your AI crawler controls are actually set to. The new default applies automatically to free-tier accounts and to any new site. Whatever you find, it should be a decision you made rather than one you inherited.
What the Cloudflare Default Actually Does
Precision matters here, because the headline coverage was loose about the scope.
| Element | Detail |
|---|---|
| Effective date | September 15, 2026 |
| What is blocked | Mixed-use crawlers — bots blending search indexing with AI training or agent use |
| Where | Only pages carrying advertising |
| Who it applies to | New Cloudflare customers, new sites from existing customers, and all existing free-tier accounts |
| Who it does not touch | Existing paid customers who have already configured their own bot settings |
| Can you override it | Yes, manually, in your bot control settings |
| Stated intent | Pressure AI operators into separating search crawlers from training and agent crawlers |
Cloudflare has also restructured its controls to divide crawlers into three categories — Search, Agent, and Training — so publishers can block or charge each independently. That taxonomy mirrors the one that should already be governing your robots.txt: the job a bot does, not the company that runs it, is what the rule should follow.
The Googlebot problem, again
The awkward part is that Googlebot is itself a mixed-use crawler. It powers Google Search and feeds AI Overviews through the same token, which is precisely why Cloudflare is pressuring operators to split their bots. Any default that catches mixed-use crawlers catches Googlebot, and for an ecommerce brand losing Googlebot is not a trade-off worth entertaining. Apple, Google and Microsoft each offer AI opt-out mechanisms that may let their crawlers stay in scope, but if you carry ads and you are on a default configuration, this is worth verifying rather than assuming.
The Crawl-to-Referral Ratios Behind It
The numbers driving publisher anger are worth understanding even if you land on the allow side, because they are the strongest version of the opposing argument.
Reporting on the Cloudflare announcement put Anthropic's crawler at roughly 38,000 page fetches for every one referral visit sent back to a publisher. OpenAI's ratio came in nearer 1,091 crawls per referral. Both are enormous compared with traditional search, where the implicit bargain was that indexing bought you traffic.
Read the metric carefully
Those figures measure crawls per referral visit, not per appearance in an answer. Being named inside an AI response has real brand value that produces no click and therefore appears nowhere in that ratio. For a brand, an uncited-but-recommended mention can still generate a sale later. For an ad-supported publisher, it monetizes at exactly zero.
That distinction is the entire fight, and it is also why the correct answer differs so sharply between publishers and ecommerce brands. If your revenue comes from impressions on your own pages, a visit is the only thing that pays. If your revenue comes from selling a product, being recommended is the thing that pays and the visit is incidental.
An appearance in an AI answer does something for a brand and nothing for an ad stack. That single asymmetry explains why publishers and ecommerce brands should reach opposite conclusions from identical data.
The Case For Blocking, Steelmanned
Presented as its strongest advocates would put it, without hedging.
- You are subsidizing a competitor. Every page an AI system ingests improves a product that increasingly answers the questions your content was written to answer, without sending anyone to you. You are paying bandwidth to reduce your own future traffic.
- The exchange is not reciprocal. Search crawling came with an implicit bargain: index my content, send me visitors. At thousands of crawls per referral, that bargain has collapsed and continuing to honor it is habit rather than strategy.
- Blocking creates negotiating leverage. Content that cannot be taken freely can be licensed. Several large publishers have converted blocks into paid agreements, and no one negotiates from a position of already having given the thing away.
- The infrastructure cost is real. More than half of AI crawl traffic re-fetches unchanged pages. You are paying for compute and bandwidth that produces nothing for anyone.
- Consent should be affirmative. A default of open access to any commercial actor who shows up with a user-agent string is a strange norm for any other asset a business owns.
- Blocking is reversible. You can unblock tomorrow. Content already absorbed into a training run cannot be recalled.
That last point is the sharpest one, and it deserves more weight than it usually gets. The asymmetry between a reversible block and an irreversible ingestion is a genuine argument for caution rather than openness.
The Case For Allowing, Steelmanned
Same treatment for the other side.
- Absence is not neutrality. If you block, the question still gets answered — using your competitor's content. You have not withheld yourself from the conversation, you have removed yourself from the shortlist.
- Your content is marketing, not inventory. For an ecommerce brand, a blog post exists to sell a product. It has no standalone licensing value. Withholding it protects an asset that does not exist while forfeiting distribution that does.
- AI referral traffic converts far better. The visitor arrives having been recommended rather than having found you in a list of ten. That quality premium is well documented and it applies to a channel you would be opting out of entirely.
- Licensing leverage is concentrated, not distributed. Major publishers with distinctive archives can negotiate. A brand with 200 posts about product selection is not going to be offered a deal, so the leverage argument is borrowing a benefit that will not arrive.
- Citations compound. Entity recognition and topical authority build slowly and non-linearly. Blocking during the establishment phase means competitors become the corroborated default, and displacing an established entity is far harder than establishing one.
- The retrieval layer is separable. You do not have to choose between all-in and all-out. Training and retrieval run on distinct user-agents, so the actual choice space is wider than the debate implies.
The Six-Question Decision Framework
Work these in order. The first three usually settle it.
If people pay for access to your content, blocking is defensible. If your content exists to sell something else, blocking protects an asset that does not exist.
Ad impressions require a pageview and die without it. Product sales survive an uncited recommendation. This single answer flips the calculation.
Would an AI company notice your absence? Distinctive archives and proprietary data have leverage. A well-executed brand blog does not, and pretending otherwise costs you distribution for nothing.
Measure before deciding. If AI referrals are already meaningful, blocking has a known price. If they are zero, find out whether that is because you are blocked before concluding the channel is worthless.
If competitors are open and you block, you hand them the category. If the whole category blocks, models answer generically and nobody wins. Check rather than assume.
If you expect AI-mediated discovery to grow, entity establishment now is cheap and later is expensive. If you expect it to plateau, the urgency drops considerably.
How the answers usually resolve
If you answered no to question one and product sales to question two, you are an ecommerce brand and the answer is almost certainly allow retrieval, decide training on policy. If you answered yes to one and ad impressions to two, you are a publisher and blocking or charging is defensible. Question three is the honesty check: most businesses that believe they have licensing leverage do not.
Publishers vs Ecommerce Brands
Most of the writing on this question is by publishers, for publishers, and it is being read by ecommerce operators as though it applies to them. It largely does not.
| Dimension | Publisher | Ecommerce Brand |
|---|---|---|
| What content is | The product | Marketing for the product |
| How a visit pays | Ad impression, requires the pageview | Product sale, survives an uncited mention |
| Uncited recommendation | Worth nothing | Worth a great deal |
| Licensing leverage | Real for large archives | Effectively none |
| Cost of invisibility | Lost impressions | Lost consideration-set membership |
| Reasonable default | Block or charge | Allow retrieval, decide training on policy |
The confusion is understandable. Publishers write more, write faster, and have a direct interest in the outcome. But an ecommerce brand adopting a publisher's posture is copying a strategy built on an economics it does not share.
The Middle Path Most Brands Should Take
Because training and retrieval run on separate user-agents, the practical choice is not binary. The configuration most ecommerce brands should land on:
- Allow every retrieval crawler unconditionally. These are how citations happen. Blocking them removes you from AI answers with no offsetting benefit.
- Allow every user-triggered fetcher. Blocking these only breaks the moment a real person tries to hand your page to their assistant.
- Decide training collectors on policy. Most brands leave them open and lose nothing. Blocking them is defensible and costs little. This is the one genuine judgment call.
- Block the non-compliant. Bytespider and similar high-volume, low-return crawlers, enforced at the server rather than in robots.txt.
- Protect the paths that should never be crawled. Cart, checkout, account, admin. This has always been true and has nothing to do with AI.
- Verify your CDN agrees. A permissive robots.txt means nothing if Cloudflare is challenging crawlers above it. After September 15 this check moves from good practice to necessary.
The implementation detail for each of those layers is covered in our AI crawler list for ecommerce, and the diagnostic for finding what is silently blocking you is in the AI crawler technical audit.
Check your config before September 15
We will look at your robots.txt, your Cloudflare bot settings, and your current AI referral share together, and tell you what the default change will do to you.
Book a Strategy Call →The Ecom Profit Box
Eleven playbooks on listings, conversion, images, and email. Built for operators, no fluff, no email sequence.
Grab It Free →Pay Per Use and Whether To Engage
Cloudflare has evolved its Pay Per Crawl marketplace into Pay Per Use, which compensates publishers when their content creates value inside an AI answer rather than merely when a bot fetches the page. That is a meaningful design change — it addresses the central complaint about flat per-crawl pricing, which was that fetches and value are only loosely related.
Cloudflare has also announced partnerships intended to compensate publishers when content appears in search results or agent queries, and a business insights dashboard surfacing referral volumes and crawl-to-referral ratios per crawler — data specifically useful for licensing negotiations.
Should an ecommerce brand engage with it?
Probably not yet, and the reasoning is unglamorous. Monetization frameworks reward content with standalone value: original research, distinctive archives, proprietary data. A brand blog optimized to sell products has little of that. The realistic outcome of engaging is administrative overhead in exchange for a payment that rounds to nothing, while the citation value you were already getting for free is what actually moves your revenue.
Where it becomes genuinely worth watching is if you publish original data nobody else has — industry benchmarks, proprietary survey results, category research. That content has real standalone value and the calculation changes.
Ignore the payments and pay attention to the reporting. Per-crawler referral volumes and crawl-to-referral ratios are exactly the data most brands currently lack. Knowing which AI systems consume your content and which actually send anyone back is useful whether or not you ever charge for access.
What Actually Happens After You Block
A realistic timeline, because the consequences are slower and quieter than people expect — which is exactly what makes them easy to misattribute.
| Window | What you observe | What is happening |
|---|---|---|
| Days 1–7 | Nothing | Indexes still hold your content; answers unchanged |
| Weeks 2–6 | Still mostly nothing | Cached content ages out gradually rather than dropping |
| Months 2–4 | Citations thin out | Fresh content never enters the index; older mentions persist |
| Months 4–8 | Competitors named instead | Your slot in the answer gets filled by whoever stayed open |
| Months 8+ | Absent from the category | Entity association decays; you are no longer a default answer |
| After unblocking | Slow return | Re-establishing an entity takes materially longer than losing it |
The dangerous feature of this curve is the flat start. Nothing visible happens for six weeks, which reads as confirmation the block was costless. The cost arrives in month four, by which point almost nobody connects it to a robots.txt change made a quarter earlier. Our analysis of how AI citations compound over twelve months covers the same curve running in the other direction.
The Risk Nobody Is Pricing: Other People's Blocks
Here is the second-order effect that gets almost no attention, and it affects you even if you never touch a setting.
Your citations do not only come from your own pages. They come from roundups that mention you, forum threads that recommend you, review sites that cover you, and industry publications that quote you. When those sites block AI crawlers — or get blocked by a CDN default they never configured — the pages that were carrying your brand into AI answers disappear from the retrievable corpus.
You did nothing. Your configuration is unchanged. Your visibility drops anyway.
What this implies practically
- Owned content matters more than it used to. Pages you control cannot be blocked by someone else's default.
- Diversify where your mentions live. If your citation profile depends heavily on two or three ad-supported publishers, that is now a concentration risk with a date on it.
- Re-baseline after mid-September. If your AI visibility moves in October, third-party blocks are a candidate explanation before anything you did.
- Platforms with different economics are more durable. Community sites, forums, and non-ad-supported sources are less exposed to this particular change.
This is the strongest available argument for building a deep owned-content base rather than relying on earned mentions, and it is covered further in our piece on building a compounding content moat.
Decision Matrix by Business Model
The summary answer for each situation, with the reasoning compressed.
| Business model | Training bots | Retrieval bots | Why |
|---|---|---|---|
| Ecommerce brand | Allow, or block on policy | Allow | Content is marketing; citations sell product |
| DTC with content engine | Allow | Allow | Entity establishment is the whole play |
| Local service business | Allow | Allow | Thin corpus means every source counts |
| SaaS / B2B | Policy choice | Allow | Docs and comparisons drive high-intent discovery |
| Ad-supported publisher | Block or charge | Evaluate carefully | Pageview is the unit of revenue |
| Subscription publisher | Block | Allow teasers only | Content is the product being sold |
| Original research / data | Block | Allow with attribution | Genuine licensing leverage exists |
| Marketplace seller | Allow | Allow | Product discovery happens off your domain anyway |
The three things to do before September 15
- Check your CDN bot settings. Find out what is configured today. If you are on a free tier or a recently created site, the default is about to change under you.
- Measure your current AI referral share. You cannot evaluate the cost of a block without knowing what you are giving up. The AI visibility audit guide covers the diagnostic, and several tracking tools automate the sampling.
- Write the decision down. Whatever you choose, record why, so that in four months when something moves you can tell whether this was the cause.
The worst outcome is not blocking or allowing. It is discovering in December that a setting you never saw made the decision for you in September.
The Short Version
- From September 15, 2026, Cloudflare's defaults block mixed-use crawlers from ad-supported pages, applying automatically to new customers, new sites, and all existing free-tier accounts.
- By June 2026 training crawlers were roughly 50.6% of AI bot traffic on Cloudflare's network while search bots had fallen to about 10.7%, and over half of AI crawl traffic re-fetches unchanged pages.
- Crawl-to-referral ratios reported around 38,000 to 1 for Anthropic and 1,091 to 1 for OpenAI measure referral visits, not answer appearances — which is why publishers and brands should reach opposite conclusions.
- For most ecommerce brands the answer is allow retrieval unconditionally and decide training on policy, because content is marketing rather than inventory and you have no real licensing leverage.
- Blocking has a flat consequence curve: nothing visible for six weeks, citations thinning at months two to four, competitors occupying your slot by months four to eight.
- The unpriced risk is other people's blocks. When sites that mention you go dark to crawlers, your visibility falls without you changing anything.
- Before September 15: check your CDN bot settings, measure your current AI referral share, and write down the decision and the reasoning.
External Sources Cited in This Article
- TechCrunch — Cloudflare's September 15 default change and Pay Per Use
- Cloudflare Blog — Pay per crawl and crawler authentication
- Cloudflare Docs — Bot management and AI crawler controls
- OpenAI — Crawler documentation and user-agent separation
- Anthropic — Web crawler documentation and site owner controls
- Google Search Central — Crawler and user-agent overview

