How Cloudflare's New Default Could Silently Block AI From Seeing You
Most businesses that lose AI visibility lose it for predictable reasons: inconsistent NAP data, missing schema, no directory presence, star ratings below the ChatGPT exclusion floor. These are problems you can see and fix.
There is a less visible problem that is about to become much more common -- and most businesses affected by it will not know it happened.
What Cloudflare launched on July 1
On July 1, 2026, Cloudflare launched the waitlist for its Monetization Gateway -- a system that lets any publisher behind Cloudflare charge AI companies for content access, settled in stablecoins at a minimum of $0.01 per successful retrieval. Publishers can set three independent policies per crawler type: Allow (free access), Charge (pay to crawl), or Block (no access at any price).
That part is opt-in. The part that isn't opt-in arrives in September.
Starting September 15, 2026, every new domain joining Cloudflare will have Training and Agent bots blocked by default on ad-carrying pages. Not opt-in blocked -- just blocked, unless the site owner actively configures otherwise.
In our 2026-07-21 investigation of this mechanism (Scout session 83), we documented Cloudflare's three-category crawler taxonomy:
- **Search bots** (Googlebot, Bingbot, DuckDuckGo) -- always allowed free under the default policy - **Training bots** (GPTBot, ClaudeBot, Google-Extended) -- used by AI companies to build training datasets; target of the monetization gateway - **Agent bots** -- autonomous AI agents executing tasks; blocked by default on ad-carrying pages for new Cloudflare domains after September 15
A roofing contractor or dental practice that launches a website on a Cloudflare-backed hosting provider after September 15 will have Training and Agent bots blocked from day one. The owner will not know this happened. The site will otherwise look and function normally.
The question that matters most for live citations
Here's where the picture gets more complicated -- and where we want to be honest about what we know versus what's still unclear.
The bots that affect AI citation in practice are not the same as training bots. When Perplexity answers a query, it runs a live web search. When ChatGPT with browsing enabled answers a query, it retrieves through Bing. These are real-time retrieval operations, not training data collection.
The critical question our 2026-07-21 investigation flagged as unconfirmed: are Perplexity's live search crawlers and ChatGPT's Bing-mediated retrievals classified as "Search" or "Agent" bots under Cloudflare's taxonomy? If they fall under "Search," the September 15 default leaves them unaffected. If they fall under "Agent," a new SMB on Cloudflare could be invisible to live AI queries from day one.
Cloudflare has not published a complete crawler classification list. What's confirmed is that Training and Agent blocking affects data collection crawlers. Whether it catches real-time retrieval -- the mechanism that matters most for immediate AI citation -- depends on how each platform's live search crawler is classified. We're monitoring this and will update methodology when that's confirmed.
This uncertainty is exactly why AI crawler access needs to be a pre-flight check, not an afterthought.
Who this compounds the most
Our home-services citation research (Scout sessions 45 through 82, most recently updated 2026-07-20) found that 87% of independent HVAC and plumbing contractors already have effectively zero AI citation share in their own market -- even those with hundreds of five-star Google reviews. The 1.2% ChatGPT citation rate for local contractor locations holds consistently across roofing, plumbing, and HVAC verticals.
These businesses start from a structural disadvantage. The core problem is entity infrastructure -- schema, directory presence, NAP consistency -- not content volume or review counts. Google reviews don't translate to AI citations because they're rendered via JavaScript and largely inaccessible to AI crawlers. Yelp, BBB, and static directory listings are AI-readable; your Google review count is not.
A new contractor who launches a site on Cloudflare after September 15 now has a potential additional blocker on top of everything else. Schema work, directory listings, and content fixes produce zero measurable impact if the crawlers implementing them can't access the site.
This is not theoretical. In our current audit pipeline, if a site has Training bots blocked, the website-level fix recommendations still generate -- we don't check AI bot access before recommending schema updates. That's a methodology gap we're closing, per the pre-flight recommendation that came out of this investigation.
The check that should come first
Scout's 2026-07-21 methodology recommendation frames this as Priority 0: before any schema, content, or directory work begins, verify that AI crawlers can actually reach the website.
Two specific checks:
**robots.txt** -- the familiar one. Go to `yourdomain.com/robots.txt`. Look for `User-agent: GPTBot`, `User-agent: PerplexityBot`, or `User-agent: ClaudeBot` with `Disallow: /`. Also look for a blanket `User-agent: * / Disallow: /` rule that blocks every compliant crawler. This is the robots.txt problem we've documented before; it applies here too.
**Cloudflare configuration** -- the newer one. If the site runs behind Cloudflare, check the bot management settings. After September 15, new domains will show Training and Agent bots as blocked by default unless the owner has changed the policy. Existing sites that joined Cloudflare before September 15 keep the Allow default -- but it's worth confirming, especially if the site was recently rebuilt or migrated.
One important distinction: directory presence recommendations are not affected by AI bot blocking on the client's own website. Directories make their own crawler policies independently. If your Yelp listing, GBP, and BBB profile are strong, those signals remain available to AI systems regardless of what your website does with robots. This is part of why directory presence isn't just good AEO strategy -- it's the part of your AI citation infrastructure that doesn't depend on your website's crawler settings at all.
The honest assessment
Cloudflare's gateway is in opt-in mode now and stays there for existing sites. The September 15 change only affects new domains. If your website has been live for months or years and you haven't explicitly configured Cloudflare to block AI bots, you're unaffected.
But if you're launching a new site in the next few months -- or advising a client who just went live on a Cloudflare-backed host -- this is worth checking in the first five minutes of any AEO review. A blocked AI crawler produces the same zero-visibility output as a business that doesn't exist yet. No amount of schema work changes that outcome.
Signal Check flags robots.txt misconfigurations in its technical section today. The Cloudflare-specific policy check is being added ahead of the September 15 default change, so it gets surfaced automatically when it starts affecting new sites rather than surfacing six months later when a client asks why their visibility score isn't moving.
See how your business scores on AI platforms.
Check your score — free