On September 15, Cloudflare began applying its blocking defaults to the three crawlers that publishers actually care about: Applebot, Bingbot and Googlebot. It paired that with a new setting called Disallow AI Training, which lets a site refuse to feed AI models while staying fully indexed in search. The coverage was close to unanimous, and close to identical: the old trade-off between search visibility and training consent is finally dead.
What almost nobody looked at is the mechanism deciding which crawlers still reach your pages. Cloudflare invented a status called Accountable, and the crawlers holding it keep crawling. It is granted against four published requirements, and the largest operators hold it partly on features they have not built yet. Microsoft’s gap is the widest. The requirement sitting at the top of Cloudflare’s own list, a robots.txt preference for refusing training, is targeted at Bingbot for early 2027. Bingbot carries the Accountable label today.
What the Label Gates
The four requirements are specific. An operator must offer a way to opt out of AI training through robots.txt or a comparable standard. It must offer a way to opt out of AI summaries, set with the operator now and through Cloudflare next year. It must provide URL-level visibility into which pages were made available for training, plus metrics on how that content surfaced in search. And it must guarantee that opting out of training does not degrade ordinary search rankings.
Meet them, and your crawler is mixed-use but trusted, so Cloudflare’s new Disallow AI Training setting lets it index for search while honouring a no-training preference. Fail them, and your training crawler gets blocked outright.
That is a meaningful piece of infrastructure. It is also, functionally, a licence to read the web, issued by a company that routes a large share of it and answers to no electorate on the question.
The Ledger Is Mostly Promises
Read Cloudflare’s own accounting of who has shipped what, and the picture is less settled than the headlines suggested.
Apple has shipped the robots.txt opt-out through Applebot-Extended, the nosnippet directive, and labelling for paywalled content. Its URL-level inspection tooling, requirement three, is due next year.
Google is furthest along. The robots.txt opt-out runs through Google-Extended, there is a toggle in the webmaster portal, and reporting exists. Cloudflare still lists additional URL-level transparency as arriving in the coming weeks.
Microsoft has shipped the NOARCHIVE meta tag and a Block URLs tool. Its support for a robots.txt no-training preference, the first and most basic requirement on the list, is targeted for early 2027. That is roughly sixteen months of holding a compliance badge for a control that does not exist.
A standard that admits its three biggest members on credit, then blocks everyone else on delivery, is not a standard. It is a seating chart.
Amazon, Anthropic, Meta and OpenAI are also categorised as Accountable, on the narrower and cleaner basis that they run separate search and training crawlers. Their training-specific crawlers stay blocked when a site publishes the no-training preference, which is the outcome the setting is supposed to produce.
Why the Asymmetry Matters
The practical effect is a two-tier system sorted by size. A large operator with a roadmap gets the designation and keeps its access. A smaller AI company that has shipped nothing gets blocked, which is the correct treatment, but the two are being judged on different evidence. One is assessed on what it has built and one on what it has undertaken to build.
Cloudflare’s incentive here is not hidden. A default that blocked Googlebot from ad-supported pages would be unusable for most of its customers, because search referral traffic is still the thing paying for the content. The company needed Google, Microsoft and Apple inside the tent for the policy to function at all. Admitting them on commitments was the way to get there by September 15.
That is a defensible product decision and a weak governance one, and BTN thinks Cloudflare should stop pretending they are the same thing. If the designation is going to decide who reads the open web, it needs the machinery of an actual standard: a published compliance date per operator, a public status page showing shipped against promised, and a stated policy for what happens when a deadline slips. Right now there is no visible consequence for Microsoft missing early 2027. A commitment without a penalty is marketing.
The fix is not complicated. Cloudflare already publishes the per-operator breakdown. Turning that into a dated scorecard with revocation attached would cost it very little and would convert a courtesy into a contract.
The Payment Layer Is the Real Prize
Sitting behind the crawler rules is the business Cloudflare is actually building. Pay Per Crawl, the metered experiment it launched in private beta in 2025, is becoming Pay Per Use, with Ceramic.ai and You.com as launch partners. The shift is from charging for a fetch to charging when content creates value, which means a publisher gets paid when its material shows up in an AI answer.
Publishers should read that change carefully. A fetch is an event you can see in your own logs. An appearance inside someone else’s generated answer is an event only the AI company can count. Moving the billable unit from the first to the second hands the meter to the party writing the cheque, and requirement three, the URL-level visibility that would let a publisher audit any of it, is precisely the requirement most operators have not delivered.
That ordering is the thing to watch. The transparency tooling should have landed before the payment model that depends on it, not next year.
What Changes Now
For most site owners the immediate work is small. Existing Cloudflare customers have had their settings migrated automatically, and new domains are being preset according to whether they carry ads. Publishers who want search traffic without training exposure now have a switch that did not exist a week ago, and it is worth using.
The larger shift is that the terms of access to the web are now being set in a product changelog rather than in a statute, and the first three companies through the gate were graded generously. Cloudflare has taken on a job nobody else was doing. It should now accept the part of that job it has so far skipped, which is holding the largest operators to the same evidentiary standard it applies to everyone else.