Cloudflare announced on 16 September 2026 that websites can now refuse AI training while staying fully in search results. The new “Disallow AI Training” setting ends a tradeoff publishers have complained about for three years. Alongside it, Cloudflare introduced an “Accountable” designation for well-behaved crawlers — and Apple, Google, and Microsoft have already earned it.
The numbers explain the urgency. Mixed-use crawlers, which collect content for search and AI training in a single bot, now account for 36.6% of verified crawler traffic on Cloudflare’s network. That is the single largest category. Meanwhile, fewer than 1% of website owners block search crawlers, while 17% restrict AI training. Until now, refusing the training meant risking the search.
The timing is not casual. Eleven weeks ago, Cloudflare issued the AI industry an ultimatum with a September 15 deadline. The terms: separate search from training, give site owners real controls, or get blocked from ad-carrying pages. This release is the deadline arriving — and the big three search companies blinking first.
The Cloudflare AI Crawler Controls, Decoded
The announcement replaces Cloudflare’s old “Block AI Bots” switch with three independent toggles: one each for search, AI training, and AI agents. A new customer’s recommended settings now depend on the business model. Sites that carry ads get search on, training off, and agents blocked on ad pages. Everyone else gets everything allowed.
The Accountable bar
Four criteria define the Accountable designation. An operator must let site owners opt out of AI training through robots.txt or a comparable standard. It must allow opt-outs from AI-generated search summaries. It must provide URL-level visibility into how content gets used. And it must publicly confirm that opting out of training will not hurt search rankings.
Under “Disallow AI Training,” Accountable mixed-use crawlers — Applebot, Bingbot, Googlebot — keep crawling for search while honoring the no-training preference. Training-only crawlers from Amazon, Anthropic, Meta and OpenAI get blocked outright, since blocking them never affected search anyway.
What replaced what
“Managed Robots.txt” becomes “Bot Preference Sync,” which pushes a site’s crawling preferences across all supported crawlers at once. The changes apply automatically to new Cloudflare sites and to free-tier customers who never touched their settings — defaults that Cloudflare first threatened in July.
Why the Controls Arrived Now
This is the third act of a campaign that began in July 2025, on what Cloudflare called Content Independence Day. Act one defaulted new Cloudflare domains to blocking AI crawlers. Act two, on July 1, 2026, sorted all AI traffic into Search, Agent and Training categories and set the September 15 deadline.
The backdrop moved too
Bots now drive more than half of all web requests — a milestone that arrived earlier than forecast. Cloudflare’s network alone sees over 50 billion AI-crawler requests per day. Apple confirmed in June that Applebot data feeds its foundation models. Microsoft’s robots.txt-based no-training mechanism, meanwhile, is only “targeted for early 2027” — today it offers a NOARCHIVE meta tag as a workaround.
Cloudflare AI Crawler Controls and the Competition
Cloudflare is not the only gate, and rivals point out the limits of the Cloudflare AI crawler controls.
Fastly plays hardball
Fastly hard-blocks GPTBot on 61.5% of observed requests, versus Cloudflare’s 16.9% hard blocks plus a 24.7% challenge layer, according to a July 2026 measurement study. Moreover, Fastly has also criticized Cloudflare directly. An August 2025 blog post argued Cloudflare’s early controls exempted Google and Apple. Fastly instead pointed to its TollBit partnership, which lets publishers charge bots instead of banning them. What Fastly lacks is Cloudflare’s designation system and default-block posture.
Akamai and CloudFront
Akamai offers sophisticated general bot detection but no AI-specific product line, no content signals and no monetization mechanism. AWS CloudFront serves 78.5% of GPTBot traffic without intervention, per the same study. The open-standards alternatives — the IETF’s draft “ai-prefs” specification, Cloudflare’s own Web Bot Auth proposal — remain works in progress.
What the Public Data Shows
Independent measurement suggests the celebration should stay modest. A July 2026 study of 10,894 top-web domains found only 26% behind Cloudflare’s proxy. Add the ad-carrying condition, and the September 15 defaults touch 8.5% of the top web — meaning 91.5% sits outside Cloudflare’s enforcement entirely.
The enforcement gap is real even on-network. Of sites that ban GPTBot in robots.txt, 39.5% still serve it on the wire, the study found. Google’s AI Mode also cites sites that block AI crawlers more than three times as often as the baseline. Disallowing training does not remove you from AI answers.
The company behind the controls
Cloudflare’s Q2 2026 results, reported August 6, show revenue of $696.1 million, up 36% year over year. Non-GAAP net income reached $107.8 million, or $0.29 per share. The company holds $4.16 billion in cash and securities and guided full-year revenue to roughly $2.87 billion. CEO Matthew Prince told investors the web is being “rewritten for machine-to-machine traffic” and that Cloudflare sits at the center of that shift. A GAAP operating loss of $205.7 million, inflated by a $150.7 million restructuring charge, barely dented the narrative.
New vs Repackaged: What Actually Changed
New — the settings. Disallow AI Training, the three-way control split, business-model-based defaults and Bot Preference Sync all went live this week. These are real product changes, not positioning.
Improved — the big three’s posture. Nothing Google and Apple offer here is technically new; Google-Extended and Applebot-Extended have existed since 2023 and 2024. What improved is the commitment to honor a single no-training preference inside their mixed crawlers, on a published timeline.
Repackaged — the standards story. The IETF ai-prefs participation and the Content Independence Day framing recycle a year of Cloudflare blog posts.
Unclear — verification. The release never explains how Cloudflare audits compliance, what happens when an Accountable crawler breaks a commitment, or who decides.
The Questions the Press Release Doesn’t Answer
Who referees the referee? Cloudflare writes the criteria, grades the crawlers, operates the blocking infrastructure and sells Pay Per Crawl monetization on top of it. A private company now sets de facto web policy across a fifth of the internet — an unelected standards body with a marketplace attached.
What about the other 91%? The controls bind only Cloudflare’s network. Everywhere else, robots.txt remains a voluntary courtesy that roughly two in five GPTBot bans fail to enforce.
What did Microsoft actually concede? A robots.txt training opt-out “targeted for early 2027” — a promise, not a product. Today’s option is a NOARCHIVE tag designed for archival, not training.
What about AI summaries? The release admits this fight is next. The current opt-out is a yes-or-no toggle per operator, with granular controls promised “by early next year.”
Who Should Act on the Cloudflare AI Crawler Controls
If your site runs on Cloudflare, act today. Your defaults may have changed on September 15. Review the three toggles deliberately — search, training, agents — and confirm Bot Preference Sync matches your intent. The tradeoff between visibility and control is now yours to set, not the crawler’s.
If you publish elsewhere, read this as a roadmap, not a solution. Your robots.txt is still voluntary. Edge enforcement is the only mechanism that actually blocks a non-compliant bot, and Fastly, Akamai and origin-level tools each handle it differently.
If you operate an AI crawler, the mixed-use model is now officially a liability under the Cloudflare AI crawler controls. Separate your search and training crawlers, publish opt-out mechanisms, or accept blocked access to a growing share of ad-funded content. The Accountable list is public on Cloudflare Radar; absence from it is now a commercial disadvantage.
If you watch web governance, the question is not whether the Cloudflare AI crawler controls work. It is whether one company should write the rules. Can the IETF’s ai-prefs effort produce an open standard before Cloudflare’s version hardens into one?

Editor’s Note
This article draws on Cloudflare’s 16 September 2026 press release and its July 1, 2026 announcement; product details, designation criteria and network statistics are company-reported. It also uses Cloudflare’s Q2 2026 investor materials and a July 2026 measurement study by HasData covering 10,894 domains. Fastly’s published blog positions and reporting from TechCrunch and The Register complete the sourcing. Independently verified: the announcement timeline, the July deadline, the three-way control structure, competitor positions and Cloudflare’s financial results. Not verified: the 36.6% mixed-use crawler share beyond Cloudflare’s own network, the mechanics of Accountable compliance auditing, and whether designated operators’ commitments hold past their stated timelines.

