If a site you run sits behind Cloudflare, a setting you may never have opened may have changed meaning. Cloudflare now sorts AI crawlers by what they do, and since 15 September 2026 it applies new defaults to pages that show ads. The awkward part is that Google's main crawler does more than one thing, and Cloudflare's rule for that case is the strictest one available. Nobody has published evidence that this has cost a site its Google traffic. It is worth twenty minutes to confirm it has not cost yours.
What changed on 15 September
Cloudflare announced the change on 1 July. It now groups AI traffic into three behaviours. Search is "any behavior that collects or indexes your content, so it can answer questions about it later". Agent is "automated behavior that is acting, usually in real time, on a person's behalf, to get something done right now". Training is "a crawler taking your content to train or fine-tune a model".
Each behaviour can be managed separately. From 15 September the default, in Cloudflare's developer changelog, is that "Bots classified as Training or as Agent are blocked on pages that display ads, while Search remains allowed." Ad-funded pages earn from visitors, many of whom arrive from search, which is why the Googlebot question matters.
Who is affected: Cloudflare's two pages disagree
Cloudflare's blog post and press release, both dated 1 July, describe the scope differently.
| Cloudflare page | Who gets the new defaults |
|---|---|
| Blog post and developer changelog | "all new domains onboarding to Cloudflare". The changelog says all customers could opt out "at any time before September 15". |
| Press release | "new customers and for new sites for existing customers", plus "all existing free customers that have not changed their settings by September 15, 2026". |
The press release is the wider reading, and it covers the sites most likely to have been set up once and left alone: small business sites on the Free plan. Do not work out from the plan name whether you are affected. Open the dashboard and read what the settings say today.
Why a search crawler gets the strictest rule
Cloudflare's blog is explicit about crawlers that do both: "Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service)." The changelog adds that "multi-purpose crawlers that combine Search and Training will be affected by the new defaults to block Training." The press release goes further for new customers and new sites: "Mixed crawlers that do not give site owners the ability to choose between search, agent use, and training, will be blocked on all pages with ads."
Google's own training control is Google-Extended, and Google's crawler documentation says it "doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings". A firewall looking at one request from Googlebot cannot tell whether that page will feed Search, Gemini training or both. Cloudflare's view, in its press release, is that "mixed crawler bots disadvantage both site owners and transparent AI companies".
The audit, step by step
1. Read the dashboard settings
In the Cloudflare dashboard, select the domain and open Security settings, where Cloudflare's changelog points for these options. Write down the setting for Search, Agent and Training. Check whether the legacy Block AI bots toggle is on, since Cloudflare names it as a second route to the same result. Then open AI Crawl Control, whose Crawlers tab shows whether each crawler is allowed or blocked. Note whether the site shows ads: the new defaults act on ad pages, but Cloudflare's blog describes a hand-set Training block stopping Googlebot with no mention of ads.
2. Check robots.txt, and remember what it cannot tell you
Open yourdomain.com/robots.txt. Confirm nothing disallows Googlebot from pages you want indexed, and note whether Cloudflare is managing the file, which adds Content Signals lines. Robots.txt is a request that well-behaved crawlers follow. A Cloudflare block is enforced at the edge before the request reaches your server. A clean robots.txt proves nothing about whether Googlebot is being turned away.
3. Look for Googlebot in the logs
Filter your logs for the Googlebot user agent and compare the weeks before and after 15 September. Verify the hits are real, since the user agent is easy to fake: Google says genuine crawlers resolve to googlebot.com hostnames and crawl from the ranges in its published IP lists. If Cloudflare blocks the request, it never reaches your origin, so server logs show a quiet drop in Googlebot visits, not a run of errors. On the Cloudflare side, the Metrics tab in AI Crawl Control can be filtered by crawler to show how its requests are being treated.
4. Read Search Console crawl stats
In Search Console, open Settings, then Crawl stats. The report only appears for root-level properties, so use a Domain property or a root URL-prefix property. Check three things:
- Total crawl requests, for a step down around 15 September that did not reverse.
- Crawl responses, for a rise in anything other than OK responses.
- Googlebot type, to see whether the smartphone crawler, which Google uses to crawl content for indexing and ranking, is the one that fell.
Crawling comes before indexing, so a crawl problem can show here before it shows in impressions. Check now, not after the traffic report. If you are already watching AI Overview impressions in Search Console's AI Overviews report, annotate 15 September there too, so a crawl problem is not mistaken for click loss to AI answers.
5. Choose settings that separate search from training
If you want Google Search and do not mind Gemini training, set Training to allow. Before 15 September Cloudflare also offered an explicit opt-out in Security settings, confirming you "want no changes on Training crawlers that also crawl for Search purposes"; check whether one was recorded. The press release says the settings can be changed in the dashboard at any time. If you want Search but not Gemini training, leave Training unblocked so Googlebot gets through, and add a Google-Extended disallow to robots.txt. Google states that Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search". Bingbot and Applebot, which Cloudflare names alongside Googlebot, need the same check.
Look at the Agent setting as well. Blocking agents on ad pages also turns away AI tools fetching a page for a person in real time, which may be the wrong default for a shop or booking site that wants to be reachable by AI agents that browse client sites. If you want to be cited in AI answers, that also depends on search-type crawlers reaching you, so the same audit supports getting cited by AI search.
Sources
- https://blog.cloudflare.com/content-independence-day-ai-options/
- https://www.cloudflare.com/press/press-releases/2026/cloudflare-allows-the-agentic-internet-to-flourish-with-a-simple-philosophy-your-content-your-rules/
- https://developers.cloudflare.com/changelog/post/2026-07-01-ai-traffic-options/
- https://www.searchenginejournal.com/cloudflares-ai-crawler-rules-can-block-googlebot/581385/
- https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers
- https://support.google.com/webmasters/answer/9679690
- https://developers.cloudflare.com/ai-crawl-control/features/analyze-ai-traffic/
- https://developers.google.com/search/docs/crawling-indexing/mobile/mobile-sites-mobile-first-indexing



