Cloudflare's new AI bot defaults go live September 15: the 4 checks that save your rankings and AI citations
On September 15 Cloudflare flips new defaults: Training and Agent bots blocked on ad pages, and multipurpose crawlers like Googlebot get blocked whole if any Training rule exists. Four settings to review this week, before organic traffic and AI citations quietly disappear.
On September 15, Cloudflare changes the automated-traffic defaults for every customer: Training and Agent bots get blocked by default on pages that display ads, and Search stays allowed. The trap is multipurpose crawlers. If any rule in your account blocks Training, the whole bot gets blocked, Googlebot included. Four checks this week keep you from quietly losing organic traffic and AI citations.
This story sounds like infrastructure, which is exactly why most marketing teams will ignore it. Mistake. What Cloudflare classifies today decides who gets to read your site, and whoever cannot read it cannot recommend you on any channel. Start with what Cloudflare actually announced, then run the checklist.
What changed, exactly, and when
On July 1, Cloudflare replaced the old "Block AI Bots" toggle with a taxonomy of three behaviors (official announcement):
- Search: collects and indexes your content so it can answer questions later. The classic Googlebot and Bingbot crawl.
- Agent: acts in real time on a person's behalf. ChatGPT-User, PerplexityBot, browser-use agents driving Chrome.
- Training: absorbs your content into model training or fine-tuning.
Starting September 15, defaults for newly onboarded domains are: Training and Agent blocked on pages that display ads; Search allowed. Then there is the part almost everyone skimmed past:
multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training
Multipurpose crawlers are governed by their most restrictive behavior. A Googlebot that both indexes and feeds training data gets caught by any anti-Training rule. You do not need to touch anything on the 15th. It is enough that last year's "block AI bots" config is still sitting active in your account.
This is not hypothetical anymore. Sites are losing indexation now
In early August, a report on r/SEO previewed the problem. An admin testing the new controls set "AI Training = Block" and watched Googlebot and Bingbot start receiving HTTP 403 errors when fetching his sitemap. Disabling the Training block restored access immediately (Search Engine Journal, August 4). John Mueller asked for details via DM to investigate.
Maybe it was a bug, maybe user error. For practical purposes it does not matter: the multipurpose classification exists, it is live, and it already produces false positives against search engine crawlers. If your agency or technical team "protected" your site against AI six months ago, that decision just changed price.
Why this is a GEO problem, not just a sysadmin problem
Every generative engine optimization strategy silently assumes engines can reach your content. No access means no citation, and almost nobody audits that prerequisite before touching anything else.
Here comes the uncomfortable detail. Googlebot feeds both the classic index and the answers behind AI Overviews. An accidental Googlebot block does not just cost you traditional SEO. It costs you the surface where Google builds AI answers. One badly configured toggle takes out both channels at once.
At the other end sit real-time bots: when someone asks ChatGPT something with live browsing, or an agent drives a browser to check your pricing page, that fetch runs under a User-agent like ChatGPT-User. Block it and answers get built without you. The citation paradox is hard enough to crack with full access. Imagine cracking it without any.
The decision table: three bot types, three different decisions
| Bot type | What it does | Our recommendation | Risk if you block it |
|---|---|---|---|
| Search | Indexes to answer queries (Googlebot, Bingbot) | Allow, always | Indexation drops, AI Overviews citations vanish |
| Agent | Acts in real time (ChatGPT-User, PerplexityBot) | Allow if you want to appear in live answers and agents; consider rate limits | You disappear from answers built in real time |
| Training | Trains models (GPTBot, ClaudeBot, Bytespider) | A legitimate business decision, with a known cost | Less long-term memory of your brand inside models |
Blocking Training is a defensible stance: licensing content, protecting data, negotiating deals. What is no longer defensible is doing it carelessly and taking out the same crawler's Search function as collateral damage.
Checklist: four verifications before September 15
- Review the AI bots section in Cloudflare (Security → Bots, on every domain). Note which categories sit on Block today. If Training is blocked, check how multipurpose crawlers are treated before the new default makes it automatic.
- Confirm Googlebot and Bingbot actually get through. Look for 403s or security challenges in access logs, then run URL Inspection in Search Console. If something fails, turn off the Training block and test again. Fastest diagnosis there is.
- Compare robots.txt against your Cloudflare config. At mintec.co we keep AI bots explicitly allowed in robots.txt with matching policy at the CDN layer. Contradictions between layers are what produce surprises, so pick one source of truth.
- Decide Agent and Training with a written rationale, not by default. If your business lives on organic discovery, our advice is Search allowed, Agent allowed, Training as a conscious decision. Write the decision down: in six months nobody will remember why anything was blocked.
The contradiction nobody has resolved yet
Mueller also used the thread to remind everyone that Google ignores llms.txt and similar files. Every layer of the ecosystem pushes different norms: Cloudflare classifies bots, Google disregards third-party content signals, and AI engines do their own thing. Until that settles, your only robust play is internal coherence. Make robots.txt, headers, CDN configuration, and content say the same thing.
That is how we run our own stack at Mintec: GPTBot, ClaudeBot, PerplexityBot and friends explicitly allowed in robots.txt, mirrored at Cloudflare, with citations tracked so the decision rests on numbers instead of fear. Last time the market moved, that record saved us a week of deciding blind. It will save us again the next time.
Frequently Asked Questions
Will Cloudflare block Googlebot after September 15?
Not automatically. The new default blocks Training and Agent bots on pages that display ads while leaving Search allowed, but multipurpose crawlers like Googlebot, Bingbot, and Applebot combine search indexing with training collection. If any rule in your account blocks Training, the most restrictive rule applies to the entire bot and Googlebot stops getting through. Review your settings before September 15.
What are Search, Agent, and Training bots in Cloudflare's classification?
Search collects and indexes content so it can answer queries later (Googlebot, Bingbot). Agent acts in real time for a person: ChatGPT-User, PerplexityBot, browser-use agents. Training absorbs your content into model training. Since July, Cloudflare lets you manage each category separately, and since September 15 it applies restrictive defaults on pages that display ads.
Does blocking AI bots hurt my visibility in AI Overviews and ChatGPT?
Yes, directly. If Perplexity, ChatGPT, or agents cannot fetch your pages, they cannot cite you. Citation starts with access. On Google the risk is worse because Googlebot feeds both classic search and AI Overviews, so an accidental block removes you from everything at once. Our position: allow Search always, decide Agent case by case, think twice before blocking Training if you want to show up in AI search.



