KILLEREDGE Join The Dojo

Home / Blog / Guides

AI Search

On 15 September Cloudflare starts blocking AI crawlers by default. Here is how to check you are not walled out of the answer.

By David Sessford · 2 Sep 2026 · 11 min read

From 15 September 2026 Cloudflare blocks mixed use AI crawlers by default on ad carrying pages, for new customers, new sites and all existing free plans. Allow the search bots that cite you, such as OAI-SearchBot, Claude-SearchBot and PerplexityBot. Decide separately on training bots like GPTBot. Check your settings before the date.
Key takeaways

On 15 September 2026 a switch gets flipped that decides whether AI can read your website. You did not vote on it. Your web developer probably does not know about it. And if you are on the wrong side of it, ChatGPT stops being able to quote you and you will never see a single alert.

Cloudflare sits in front of more than 20% of the web. On 15 September its default settings start blocking mixed use AI crawlers from any page that carries ads. That applies to new Cloudflare customers, to new sites added by existing customers, and to every existing free plan customer. Paid customers with settings already in place keep them.

Read that again. Every existing free plan customer. That is a lot of small businesses who have never once opened the bot settings.

Everyone in SEO spent August arguing about AI Overviews. Meanwhile the actual door to the AI answer is being fitted with a lock, and most site owners do not know which side of it they are standing on.

What actually changes on 15 September 2026

Cloudflare announced this on 1 July 2026. The company gave AI firms a deadline: split your crawlers by purpose, or get blocked by default.

The logic is not mad. Cloudflare's own bot report says 52% of crawler requests are now for AI training, up from 22% in spring 2025. Mixed use crawlers, the ones that blend search, agent use and training into a single bot, make up more than a third of all crawler activity. Cloudflare also found that over half of AI crawl traffic is spent re fetching pages that have not changed.

So publishers are paying bandwidth to feed models that send almost nothing back. Cloudflare's answer is to make the bots declare what they are for. Bots with clear intent get through. Bots that refuse to separate search from training get blocked on monetised pages by default.

At the same time, Pay Per Crawl is turning into Pay Per Use. Instead of charging a bot to fetch your page, you get paid when your content actually shows up in an AI answer. Ceramic.ai and You.com are the first two partners. Interesting. Not yet a business model for a plumber in Gateshead.

Here is the part that matters to you: this is a default, not a law. You can change it. But defaults win, because nobody changes them. That is the whole point of a default.

The bit nobody explains: there are three kinds of AI bot

Ninety percent of the block AI bots advice floating around treats every crawler as one thing. It is not one thing. Get this distinction wrong and you will either give your content away for free or wall yourself out of the answer. There is no third mistake, but those two are bad enough.

Cloudflare's own bot reference splits them into categories. Here is what they actually do:

CategoryExample botsWhat it doesIf you block it
AI SearchOAI-SearchBot, Claude-SearchBot, PerplexityBot, ApplebotBuilds the index the assistant searches when it answers a questionYou disappear from that assistant's answers. This is the one that costs you money.
AI AssistantChatGPT-User, Claude-User, Perplexity-User, DuckAssistBot, MistralAI-UserFetches your page live, right now, because a human asked about itA real person asked the machine about you and it cannot see your site.
AI Crawler (training)GPTBot, ClaudeBot, CCBot, Bytespider, Amazonbot, meta-externalagentCollects content to train modelsYour content stays out of the training set. No direct effect on being cited today.
Search EngineGooglebot, bingbotClassic search indexing, and in Google's case AI Overviews and AI Mode tooYou leave Google or Bing. Do not do this.

Look at that table for ten seconds and the strategy writes itself. The training bots and the search bots are different bots, with different names, controllable separately. OpenAI runs GPTBot for training and OAI-SearchBot for search. Anthropic runs ClaudeBot and Claude-SearchBot. Perplexity runs PerplexityBot and Perplexity-User.

Which means you never had to choose between feeding the machine for free and vanishing. You could always have both. Almost nobody set it up.

Why this matters more than any Google update this year

In classic SEO a mistake costs you positions. You slip from three to nine, traffic halves, you notice in Search Console, you fix it.

In AI search there is no position nine. You are named in the answer or you do not exist. Blocking OAI-SearchBot does not move you down the list. It removes you from the list. And Search Console will not tell you, because Google is not the one dropping you.

That is the trap. This failure is silent. No ranking drop, no error, no email. Just a slow quiet fade out of every answer your customers are asking for, while your Google traffic looks completely normal and you carry on writing blog posts.

We drill this exact audit as the first lesson in the technical module inside the Dojo, because there is no point optimising a page for the machine that cannot fetch it. Fix the door before you paint the room.

The ten minute audit: are you walled out?

Do this today. Not next month. You have until 15 September before the defaults shift, and if you are on a free Cloudflare plan you are directly in scope.

Step 1: read your own robots.txt

Go to yoursite.com/robots.txt. Read every line. You are looking for two things.

First, any blanket User-agent: * / Disallow: / that somebody left in from a staging site. It happens more than you would believe.

Second, any AI bot rules you did not write. Plenty of WordPress security and SEO plugins added block AI bots toggles in 2024 and 2025, and plenty of people ticked them without reading which bots were on the list. If your robots.txt disallows OAI-SearchBot or PerplexityBot, you did that to yourself.

Step 2: check the CDN, not just robots.txt

This is where people get caught out. Your robots.txt can be perfect and your CDN can still be refusing the bot at the network level, before robots.txt is ever read.

In Cloudflare, that is AI Crawl Control, plus the bot settings on the zone. The controls are grouped by the same three categories: search, agent, training. Set each one deliberately. Do not inherit.

If you are not on Cloudflare, the same question applies to whatever sits in front of your site. Fastly, Akamai, Sucuri, your host's built in firewall, your WAF. Somebody there has an opinion about bots and you have never read it.

Step 3: check your firewall is not challenging the good bots

A block is loud. A challenge is quiet, and just as fatal. If Bot Fight Mode or an aggressive WAF rule serves OAI-SearchBot a CAPTCHA or a JavaScript challenge, that bot does not solve it. It gets a page of nothing, and a page of nothing is what goes in the index.

You will see this in the crawler logs as a 403 or a challenge response rather than a 200. Everything looks fine in your browser. It is not fine.

Step 4: read the logs

Cloudflare's AI Crawl Control has a crawlers view showing per bot request counts, allow and block actions, and robots.txt violations. If you are not on Cloudflare, grep your raw server logs for the user agent strings in the table above.

What you want to see: OAI-SearchBot, Claude-SearchBot and PerplexityBot hitting your money pages and getting 200s. What you do not want: zero hits, or a wall of 403s.

Zero hits from every AI search bot is not a neutral result. It means you are invisible and have been for a while.

Step 5: ask the machine

The lo fi version, and honestly the one most business owners should start with. Open ChatGPT. Type best [your trade] in [your town]. Then do it in Claude and Perplexity.

Are you named? If a competitor is named and you are not, one likely reason is sitting in your bot settings. We cover the free ways to track this properly in our guide to free AI visibility tracking, so you do not need to buy a dashboard to find out.

The setting I would run, and who should do the opposite

For the overwhelming majority of businesses reading this, the answer is boring and obvious:

Allow the search and assistant bots. Decide separately on the training bots. Never touch Googlebot or bingbot.

You want to be in the answer. Being in the answer requires being fetched. End of argument.

On training bots, it is a real choice, so here is the honest split:

Your situationAI search botsTraining botsWhy
Local service business, trades, professional servicesAllowAllowYou need every citation you can get. Nobody is licensing your boiler servicing page. Visibility beats protection.
Ecommerce and SaaSAllowAllow, usuallyProduct and comparison pages are exactly what assistants pull into buying answers. Being absent is expensive.
Publisher or media site living on ad revenueAllowBlock, or monetiseYour content is the product. This is who the Cloudflare change was built for.
Paid courses, research, proprietary dataAllowBlockYou sell the thing. Do not donate it to the training set.
Agency or consultancyAllowYour callLow stakes either way. Being cited as the expert is worth more than the IP.

Notice the first column never changes. Allow, allow, allow, allow, allow. The training column is where reasonable people differ. The search column is not a debate.

This is the bit agencies wrap in a technical AI readiness audit and invoice at four figures. It is one afternoon and a settings page. If you would rather it was just done properly and signed off, that is what Done For You is for.

What it costs

Very little, which is the annoying part, because free and quick is exactly the kind of job that never gets done.

RouteCostTime
Edit robots.txt yourself$015 minutes
Cloudflare AI Crawl Control on a free zone$020 minutes
WAF rules using bot detection IDsRequires Cloudflare Bot Management, a paid feature. See Cloudflare's plans page for current pricing. We will not quote a number we cannot verify today.1 hour
Developer does it for youTypically one to two hours of whatever your developer chargesSame day

Compare that to the cost of being absent from the answer for a year because a plugin toggle you never noticed was switched on.

The Google problem nobody wants to say out loud

Cloudflare pointed a finger directly at the world's largest search engine for running a mixed use crawler, and said it therefore has access to roughly twice as much information as other AI companies.

Google's counter is Google-Extended, a control that opts you out of training for Gemini and Vertex without affecting your inclusion in Search. Fair enough. But Googlebot itself still crawls for Search including AI Overviews and AI Mode, and you cannot separate those. Want to be in Google Search? Then you are in Google's AI features. That is the deal and there is no toggle.

Which is exactly why the AI search bots from OpenAI, Anthropic and Perplexity matter so much. Those are the ones where you actually have a lever. Pull it. If you want the full play for getting quoted in Google's own AI surfaces, that is a separate fight we have already mapped out.

Five mistakes that wall people out

  1. Blocking AI bots with a plugin toggle. You do not know what is on that list. Usually it includes the search bots. Read it or turn it off.
  2. Copying a robots.txt off a blog post. Half the AI bot lists published in 2024 predate OAI-SearchBot and Claude-SearchBot existing. You are pasting in a snapshot of a world that has moved on.
  3. Blocking the assistant bots. ChatGPT-User and Claude-User fire because a human being right now asked the machine about you. That is the warmest traffic on the internet and people block it by accident.
  4. Setting it once and never checking. Bots get renamed. New ones appear. Hosts change defaults, as Cloudflare is about to demonstrate. Diary it quarterly.
  5. Fixing the bots and nothing else. Letting the crawler in is necessary, not sufficient. If your page buries the answer 800 words down, the bot arrives, finds nothing quotable, and leaves. Read how AI actually reads a page before you celebrate.

Alternatives if you would rather be paid than read

If you are a genuine publisher with content worth licensing, blocking is a strategy, not a tantrum. Cloudflare's own report makes the case: publishers who controlled access created scarcity, scarcity created bargaining power, and that bargaining power turned into licensing deals with real attribution data behind them.

Pay Per Use is the next step. You get paid when your content appears in an answer rather than when a bot fetches it. Early days, two launch partners, but the direction is set.

Be honest with yourself about which side you are on. If you are a local business, you are not a publisher. Nobody is licensing your service pages. Your advantage is being present, not being scarce, and the way you build that presence is with mentions across the sites the models already trust, not with a wall.

Risks and what could still change

Straight with you about the limits here.

robots.txt is a request, not a lock. It authenticates nobody and stops nothing at the network level. Well behaved crawlers obey it. Badly behaved ones do not, and there are plenty. Network level rules are the only real enforcement.

The bot list moves. Every name in this article is accurate to Cloudflare's published bot reference, but operators add and rename crawlers constantly. Check the source, do not trust a blog post from six months ago, including this one.

The 15 September change is narrower than the headlines suggest. It is mixed use crawlers, on ad carrying pages, for new customers, new sites and existing free plans. If you are a paid Cloudflare customer with configured settings, nothing changes for you automatically. That is not a reason to skip the audit. It is a reason to do the audit and find out which bucket you are in.

Nobody can promise you a citation. Allowing the right bots removes a blocker. It does not create authority. That still comes from being genuinely worth quoting, which is the long game covered across our SEO work.

Where this goes next

The direction of travel is obvious once you see it. Bots will have to declare their purpose. Access will be granted or refused per purpose. Money will move based on use rather than fetch. Cloudflare says it wants mixed use crawling at zero within a year.

Which means the should I block AI question is about to stop being one question. It becomes four separate ones: train, search, agent, index. Four switches, four different answers, and a lot of businesses running all four on whatever the default was.

The winners here are not the sites with the biggest content budgets. They are the ones who spent twenty minutes making sure the door was open while everyone else argued about llms.txt. If you want the honest verdict on that particular distraction, we wrote it up already.

Open your robots.txt. Open your CDN settings. Do it before 15 September. It is twenty minutes and it decides whether the machine can see you at all.

Do it with us instead of alone

Inside the Dojo you get the audit checklist, the exact bot list kept current, the copy and paste configurations for Cloudflare and for plain robots.txt, and the rest of the AI search system: answer first page structure, entity setup, schema, and the tracker to prove it is working. No agency retainer, no twelve month contract.

Join The Dojo. $67/mo. Cancel whenever you like. Or keep guessing which bots your host is quietly turning away.

Frequently asked questions

Should I block GPTBot?

It is a genuine choice. Blocking GPTBot keeps your content out of OpenAI model training and has no effect on whether ChatGPT can cite you today, because citation runs through OAI-SearchBot. Most local and service businesses should allow both. Publishers and anyone selling proprietary content have a real reason to block training while allowing search.

Will blocking AI crawlers hurt my Google rankings?

Blocking the AI specific crawlers such as GPTBot or ClaudeBot does not affect Google Search, because Google crawls with Googlebot. What does hurt you is blocking Googlebot itself, or applying a blanket disallow that catches everything. Never block Googlebot or bingbot.

What is the difference between GPTBot and OAI-SearchBot?

They are two separate OpenAI crawlers with separate robots.txt rules. GPTBot gathers content for model training. OAI-SearchBot builds the index behind ChatGPT search, which is what decides whether you get named in an answer. ChatGPT-User is a third one that fetches your page live when a person asks about you.

Does the 15 September Cloudflare change affect my site?

It applies to new Cloudflare customers, new sites added by existing customers, and all existing free plan customers, and it targets mixed use crawlers on pages carrying ads. Paid customers with settings already configured keep those settings. Either way, check your configuration rather than assume.

Is llms.txt a substitute for getting the bot settings right?

No. llms.txt is a five minute nice to have. If your firewall or robots.txt is refusing the crawler, no file you add to the site will help, because nothing gets read. Fix access first, then worry about the extras.

How often should I recheck this?

Quarterly, and immediately after any change to your host, CDN, security plugin or SEO plugin. Those are the four things that silently rewrite bot access. Set a calendar reminder, because there is no alert when you get walled out.

Can I allow AI crawlers on some pages and block them on others?

Yes. robots.txt supports per path rules, and network level rules can target specific URL patterns. A common pattern is allowing crawlers on marketing and service pages while blocking them on gated, member only or paid content.

About the author

David Sessford. Founder of The KillerEdge, 14+ years in SEO & online marketing More about the founders →

Sources

🎁 Free. No fluff.

Grab the free AI-Search Cheat Sheet

The exact moves to make ChatGPT, Gemini and Google AI name-drop YOU. Plus the weekly intel. Free.

Instant download. No spam. Quit anytime.

Want it done for you?

Learn it in The Dojo, or have us run it. Either way, you win.

Join The Dojo. $67/mo → ← Back to the blog
It's not the battle.
It's the f*cking war.
We are the answer.