seo

Cloudflare Disallow AI Training: Keep Google, Refuse Training

Cloudflare's Block setting can stop Googlebot, while Disallow AI Training keeps Search allowed. See what each does and which one suits a small business site.

Your site runs through Cloudflare, maybe because a past developer, host or agency put it there, and you have just read a headline about Cloudflare, Googlebot and AI training. Now you are wondering whether a setting you never chose is hiding your business from Google, or from the AI assistants your customers ask. That worry is fair, and you can settle it in about ten minutes without changing anything. Cloudflare's Disallow AI Training setting exists for exactly this situation: it lets a site refuse AI training while keeping Google's crawler allowed for Search. Cloudflare's older Block setting works differently and can stop Googlebot. So, the useful questions are which setting your site has, what your live robots.txt says, and what Google Search Console shows for your key pages. The answers tell you whether you need to touch anything at all.

Key Takeaways

Check before you change

Open your live robots.txt, review the Search, Training and Agent policies in Cloudflare, and inspect a page in Search Console. Change only the rule you can show is responsible.

Disallow AI Training keeps Search open

Cloudflare added it on September 15, 2026 so a site can refuse training while Googlebot, Bingbot and Applebot stay allowed for Search. The older Block setting can still stop Googlebot.

Google-Extended is not a Search switch

Google says it does not affect a site's inclusion in Google Search and is not a ranking signal, so a Google-Extended line in your robots.txt is not what hides you from Google.

What should I check first if I think Cloudflare is hiding my site from Google?

Start with a read-only check. Nothing here changes a setting, and Google's own tools do most of the work. The one rule to keep: do not change anything until you can point to the rule that is responsible.

  1. Confirm the site is behind Cloudflare. Look at the DNS settings at your domain registrar or host to see whether your public hostname is proxied through Cloudflare. A `server: cloudflare` response header is a clue, but the DNS or dashboard view is more reliable.
  2. Open your live robots.txt. Type your web address followed by /robots.txt, using the exact `https://` address customers see. Google's guide to creating a robots.txt file explains the format.
  3. Look for three things. A `User-agent: Googlebot` group with `Disallow: /`, a wildcard `User-agent: *` group with `Disallow: /`, and Cloudflare's managed AI lines. A blanket `Disallow: /` under Googlebot or the wildcard is the one that matters for Search.
  4. Review Cloudflare's policies. In the Cloudflare dashboard, open Security Settings for your domain and look at the AI bot policies for Search, Training and Agent. Each can be Allow, Block on all pages, or Block on pages with ads, as Cloudflare's AI bot documentation describes.
  5. Ask Google what it sees. In Search Console, run URL Inspection on your homepage and one important service page. Then open Crawl Stats for response codes and recent crawl activity, and the Page Indexing report for excluded pages, noindex tags, server errors and robots.txt blocks. Google's Search Console guide walks through the reports.

If robots.txt is open, Googlebot is not blocked in Cloudflare, and Search Console inspects your pages successfully, Cloudflare is not what is hiding you. A slow month of leads then points somewhere else, such as demand, competition or the content on the page.

A closed laptop beside a clipboard holding a blank checklist and a pencil on a plain wooden desk.
The check is read-only, so nothing changes until you know which rule is responsible.

What did Cloudflare actually change, and does it affect my site?

On July 1, 2026, Cloudflare announced new options for managing AI traffic. It sorts crawlers into three groups, Search, Agent and Training, and said that from September 15 new domains would block Training and Agent crawlers on pages with ads while keeping Search allowed. Its changelog entry also warned that mixed-use crawlers such as Googlebot, Applebot and Bingbot could be blocked when a customer selected a Training block, because those bots crawl for Search and for other uses.

On September 15, 2026, Cloudflare added a setting called Disallow AI Training. In its blog post on mixed-use crawlers, Cloudflare says it is available on all plans and works by publishing the no-training preference in robots.txt. Googlebot, Bingbot and Applebot stay allowed for Search, while training-only crawlers from companies such as Amazon, Anthropic, Meta and OpenAI are refused. The older Block setting still stops crawlers outright, and Cloudflare says that includes Googlebot and Search access.

Whether it touches your site depends on which setting your domain has. Cloudflare says the September 15 defaults apply to new domains, and that existing settings carry over on their own in almost every case. "Almost every case" is a reason to look at your own dashboard rather than assume. No named, dated source shows that Cloudflare settings caused any site to lose rankings, so treat a traffic dip as something to diagnose, not as proof.

Which setting should a small business choose: Allow, Disallow AI Training, or Block?

Match the setting to what you want from search and from AI. This table sums up what each does, based on Cloudflare's documentation.

Cloudflare setting choices for training and Search
SettingWhat it doesEffect on Google SearchFits when
AllowPlaces no restriction on that group of crawlersNoneYou are happy for the content to be used for AI training
Disallow AI TrainingPublishes a no-training preference in robots.txt; keeps Googlebot, Bingbot and Applebot allowed for SearchSearch stays allowedYou want to stay findable and refuse training
BlockStops crawlers, including mixed-use crawlers such as GooglebotSearch access can be lostYou have a specific reason to keep crawlers out entirely, and you accept the Search cost
SettingAllow
What it doesPlaces no restriction on that group of crawlers
Effect on Google SearchNone
Fits whenYou are happy for the content to be used for AI training
SettingDisallow AI Training
What it doesPublishes a no-training preference in robots.txt; keeps Googlebot, Bingbot and Applebot allowed for Search
Effect on Google SearchSearch stays allowed
Fits whenYou want to stay findable and refuse training
SettingBlock
What it doesStops crawlers, including mixed-use crawlers such as Googlebot
Effect on Google SearchSearch access can be lost
Fits whenYou have a specific reason to keep crawlers out entirely, and you accept the Search cost

For a normal business site whose customers find it through Google, the safe default is to keep Search access and make a separate, deliberate choice about training and live AI agents. Allowing Search crawlers keeps you discoverable and keeps open the possibility that a page is cited as a source in an AI answer. Refusing training limits permission for models to learn from your pages. Those are different uses, and Disallow AI Training is the setting built to separate them.

One more limit worth knowing: no single setting governs every AI service. AI assistant traffic can be separate from Googlebot and from Google-Extended. If being named in AI answers matters to you, our generative engine optimization work covers that side, and it starts with the same access check above.

A padlock resting open on a wooden shelf next to a neatly closed folder, lit by soft window light.
Search access and training permission are separate choices, and you can make each on purpose.

What does Google-Extended control, and does it affect my Google ranking?

Google-Extended is a token you can put in robots.txt to say whether content Google crawls may be used to train future Gemini models and for grounding in Gemini Apps and Vertex AI. Google's crawler documentation says it "does not impact a site's inclusion in Google Search" and "is not used as a ranking signal in Google Search."

It is not a separate bot, either. Google-Extended has no HTTP user-agent string of its own. Google keeps crawling with its existing user agents, and the token records your preference. Google's AI features guidance says Googlebot is the crawler that matters for Search and AI Overviews. If you want to limit what shows in Search features, Google recommends `noindex`, `nosnippet`, `data-nosnippet` or `max-snippet`, depending on the result you want.

This matters if you switch on Cloudflare's managed robots.txt. When it is on, Cloudflare adds lines ahead of your own, including `Google-Extended Disallow: /` and lines for known AI training bots. Seeing that line in your file does not mean you are blocked from Search. The managed file is a preference signal, not enforcement. The managed robots.txt documentation says as much, and enforcement comes from tools such as AI Crawl Control or Bot Management.

What else can hide a site from Google if Cloudflare is not the cause?

When Googlebot cannot get your pages, Cloudflare is only one of several places to look. In rough order of how easy each is to check, and not a measured ranking of how often each happens:

  • An old `Disallow: /` in robots.txt, often left behind after a redesign or a staging launch.
  • A Cloudflare Block or Bot Management rule that catches mixed-use crawlers.
  • A `noindex` tag set in your website's page settings.
  • A host, DNS, SSL or server problem that keeps Googlebot from receiving pages.
  • A broad security or challenge rule that stops verified crawlers.
  • Sitemap, canonical or redirect problems that make pages harder to find.

Each has its own fix. A robots.txt block means removing or narrowing the wrong rule and testing the live file. A Cloudflare mixed-use block usually means switching the Training policy from Block to Disallow AI Training and then verifying. A `noindex` means removing the setting and confirming it in the page's HTML. Google's crawling troubleshooting guide covers the rest. Also keep in mind that robots.txt controls crawling, not removal from the index. Google points to `noindex` or password protection when the goal is to keep a page out of Search, as its robots.txt introduction explains.

A stack of blank index cards and a small brass key on a worn wooden table beside an unlit desk lamp.
A page can be blocked by robots.txt, by a noindex tag, or by a server problem, and each has a different fix.

Should I block all AI bots to be safe?

Blocking everything is not a safe first move. A broad block can stop the Search crawlers that send you customers, along with referral crawlers and verification traffic you may want. It also does not make your content untouchable, because robots.txt states a preference and a crawler that ignores it can still fetch a page. Cloudflare's Content Signals documentation describes `search`, `ai-input` and `ai-train` preferences and says they are reservations of rights under Article 4 of EU Directive 2019/790. They do not technically stop a crawler that chooses to ignore them.

If you do want firm enforcement against particular bots, that is a job for AI Crawl Control or Bot Management, set for the specific crawlers you name. Cloudflare's Pay Per Crawl, which can answer a crawler with an HTTP 402 payment-required response, is a private beta, and Cloudflare notes that WAF or Bot Management blocks override charging.

The middle route fits a site that lives on Google traffic: keep Search allowed, choose Disallow AI Training if you want to refuse training, and review the Agent policy on its own. Whichever you pick, write down the date and the setting, so the next person who opens the Cloudflare account can see what was chosen and why.

A small wooden desk sign holder with a blank card, beside a folded paper map and a pencil, on a clean table.
Write down which setting you chose and when, so the next person in the account knows why.

Is Cloudflare hiding your site from Google?

Pick an answer to begin.

1. Which Cloudflare setting keeps Googlebot allowed for Search while refusing AI training?

2. What does Google say about Google-Extended and Google Search?

3. Your traffic dipped after the news. What should you do first?

Frequently Asked Questions About cloudflare disallow ai training

What does Cloudflare's Disallow AI Training setting do?

It publishes a no-training preference in robots.txt while keeping mixed-use crawlers such as Googlebot, Bingbot and Applebot allowed for Search. Training-only crawlers are refused. Cloudflare added it on September 15, 2026, and says it is available on all plans.

Can Cloudflare block Googlebot?

Yes. Cloudflare's broader Block setting can stop mixed-use crawlers, including Googlebot, which can cost a site its Search access. Disallow AI Training is the setting designed to avoid that.

Does blocking Google-Extended remove my site from Google Search?

No. Google says Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal. It controls whether content is used for Gemini training and grounding.

Does robots.txt stop every AI crawler?

No. It states a preference, and a crawler can ignore it. Stopping a crawler by force takes enforcement such as AI Crawl Control or Bot Management.

How do I know if Cloudflare is blocking Googlebot?

Check your live robots.txt, the Search, Training and Agent policies in Cloudflare's Security Settings, and URL Inspection and Crawl Stats in Search Console. A successful inspection of your key pages is a good sign.

Did every Cloudflare customer lose Google visibility?

Cloudflare's published material does not support that. It says existing settings carry over on their own in almost every case, so look at your own account to see which setting applies.

Wrapping Up

If you are worried a Cloudflare setting is hiding your site, the check comes first: your live robots.txt, Cloudflare's Search, Training and Agent policies, and Search Console's URL Inspection and Crawl Stats. Google-Extended is not a Search switch, and a Google-Extended line in your file does not hide you from Google. The setting that can cost you Search access is a broad Block on mixed-use crawlers such as Googlebot.

For a site that lives on Google traffic, keep Search allowed and make a separate choice about training and AI agents. Disallow AI Training refuses training and leaves Googlebot in place, and you can switch it on from Security Settings for your domain. If the check turns up nothing, you can stop there, and if you did find a problem, change only the rule you can show is responsible, then retest and give Google time to recrawl. A related question you may have is why a specific page is missing from Google, and our post on pages not indexed in Search Console covers which reasons are normal and which cost you customers.

If the accounts are inherited, several layers seem to conflict, or Search Console shows a real indexing problem, Web Leveling can trace the exact rule and fix only that. Our search engine optimization work starts with the access check above, and if it shows your settings are fine, we will say so. We work with small and medium businesses across the country and overseas. Tell us what Cloudflare and Search Console are showing you, and we will help you read it.

Terms

Cloudflare and crawler words in this post

Tap a term to see what it means.

Googlebot. Google's crawler for Search. Google says it also controls crawling for AI Overviews.

Google-Extended. A robots.txt token that says whether crawled content may be used for Gemini training and grounding. It is not a separate bot and does not affect Search.

Mixed-use crawler. A crawler that gathers pages for Search and for other uses, such as Googlebot, Bingbot and Applebot.

Disallow AI Training. A Cloudflare setting, added September 15, 2026, that refuses training while keeping mixed-use crawlers allowed for Search.

Managed robots.txt. A Cloudflare option that adds AI-related directives ahead of your own robots.txt. It is a preference signal, not enforcement.

Content Signals. Preference labels in robots.txt for search, ai-input and ai-train use.

Pay Per Crawl. A private beta Cloudflare feature that can answer a crawler with an HTTP 402 payment-required response.