Cloudflare Adds a Disallow AI Training Setting That Keeps Websites Visible in Search
Website owners have faced an awkward choice for the past few years. Some of the largest crawlers on the internet collect pages for search results and for AI model training at the same time, so blocking the crawler to keep content out of training could also remove the site from search.
On September 15, 2026, Cloudflare announced a Disallow AI Training setting intended to end that trade-off. For any business that depends on organic traffic and invests in SEO services and original content, it changes how much say you have over what happens to your pages.
Below is what Cloudflare says the setting does, what changed automatically on that date, what is still to come, and what to check on your own site.
What Cloudflare announced on September 15, 2026
According to Cloudflare's blog post, the new controls are available to all customers on all plans and are configured in the Security Settings area of the dashboard. They replace a single on-or-off choice with separate controls for three kinds of automated visitor: Search, Training and Agent.
The post describes these options. Search and Agent traffic can each be set to Allow, Block, or Block on pages with ads. Training has those three options plus a fourth, Disallow AI Training, which is the new part.
The problem with mixed-use crawlers
A crawler is a program that visits pages and copies their content. Cloudflare uses the term mixed-use crawler for a single crawler that serves both search and AI training. When one visitor does both jobs, a firewall cannot let the search half in and keep the training half out.
Cloudflare's post includes its own figures on what site owners want. It says less than 1% of Cloudflare sites choose to block Search bots, while 17% of sites enable some mechanism to block training. In other words, by Cloudflare's account, nearly everyone wants to be found in search and a meaningful share do not want their content used to train models.
How the Disallow AI Training setting works
Cloudflare describes two mechanisms working together.
A published preference
For mixed-use crawlers that Cloudflare classes as accountable, a feature called Bot Preference Sync publishes a no-training preference in the site's robots.txt file, the standard text file that tells crawlers what they may access. The crawler remains allowed for search and is expected to respect the training opt-out. The post names the robots.txt tokens involved for two operators: Google-Extended for Google and Applebot-Extended for Apple.
Blocking for everyone else
The post states that when Disallow AI Training is selected, every other training crawler is blocked, including the training-only crawlers run by Amazon, Anthropic, Meta and OpenAI. Cloudflare says those four companies separate their search and training crawlers, so the training crawler can be blocked without affecting search.
The difference between the two is important. A robots.txt preference is a request that depends on the operator honoring it. A block is enforced at Cloudflare's network before the request reaches your server.
Which crawler operators are covered
Cloudflare sets out what an operator must offer to be treated as accountable: a way for site owners to opt out of AI training through robots.txt or a similar standard, a way to opt out of AI summaries, URL-level visibility into which pages were made available for training, and assurance that opting out of training will not affect traditional search results.
On individual companies, the post reports that Google has stated that disallowing Google-Extended does not impact search ranking, and that Apple has stated that disallowing training does not impact search ranking. For Microsoft, the post says the company is building a mechanism to respect a no-training preference in robots.txt at the site level, targeted for early 2027.
That last point is a future commitment, not a current capability. Until it ships, the robots.txt route described in the post is not yet in place for Microsoft's crawler.
What changed automatically and what is being retired
Cloudflare's post describes changes that took effect on September 15, 2026 without any action from customers.
Sites that had previously selected the older Block AI Bots option are migrated to Disallow AI Training for the Training control. For new domains that Cloudflare identifies as ad-supported, Training defaults to Disallow AI Training and Agents default to Block on pages with ads.
Two older features are on the way out. The post says Block AI Bots will be deprecated in favor of the separate Search, Training and Agent controls, and that Managed robots.txt will be deprecated in favor of Bot Preference Sync, with existing users migrated. We did not find a removal date for either feature in the post.
One item is announced but not delivered. Cloudflare says its goal, by early next year, is to let site owners control how much of their content appears in AI summaries with a single setting on Cloudflare instead of with each operator separately.
What this means for your business
This section is our interpretation, offered as guidance.
The decision is commercial before it is technical. A publisher, a research firm or a store with original product content may want search visibility without contributing to model training. A software company that wants AI assistants to describe its product accurately may prefer to allow training. Neither position is wrong, but it should be a deliberate choice.
Be clear about the limits. The setting covers traffic that passes through Cloudflare, so it does nothing for sites that do not use it. It does not remove content that was collected before you changed the setting. And for mixed-use crawlers it relies on the operator honoring a stated preference.
The setting concerns training, which is separate from whether your pages are quoted or summarized in AI answers. Cloudflare treats summaries as a different control that is still to come.
The Agent control also deserves thought. AI agents that browse on a person's behalf can be potential customers, so blocking them may cost you sales. If you are weighing how agents differ from simpler bots, our guide to AI agents, chatbots and workflow automation covers the distinction.
What to do next
If your site is on Cloudflare, open Security Settings and look at what Search, Training and Agent are set to today. Because existing settings were migrated, do not assume the result matches what you would choose.
Then view your live robots.txt file in a browser. If your developers or an SEO plugin maintain that file, confirm that the rules Cloudflare publishes and your own rules do not contradict each other.
Finally, record the decision and who made it. If your business licenses content, or has contracts that address AI use of that content, involve whoever handles those agreements. This is not legal advice.
Conclusion
Cloudflare's Disallow AI Training setting, announced on September 15, 2026 and available on all plans, lets a site stay open to search while publishing a no-training preference for accountable mixed-use crawlers and blocking other training crawlers. Some existing settings were migrated on that date, Microsoft's robots.txt support is targeted for early 2027, and controls for AI summaries are still a stated goal.
Entrant Technologies builds websites, web applications, mobile apps and custom software. If you would like a second opinion on how your site handles crawlers and search visibility, you can contact us.