Cloudflare has announced a website control intended to separate search-engine visibility from artificial-intelligence training, addressing crawlers that currently perform both jobs under one identity. The new Disallow AI Training setting publishes a preference while allowing participating mixed-use bots to continue indexing a site for search.

Site owners have often faced an all-or-nothing decision with these crawlers: permit access for every declared purpose or block the bot and potentially disappear from search results. Cloudflare says Apple, Google and Microsoft either already honor the new preference or have made time-bound commitments to do so. It calls operators meeting its transparency and control requirements "Accountable."

The setting uses a Disallow directive in a site's robots.txt file, but Cloudflare says enforcement involves more than publishing text. Its network identifies crawlers, classifies their behavior and can block training bots that disregard the expressed choice. The company plans to report observed operator behavior through Cloudflare Radar.

The distinction reflects different attitudes among Cloudflare customers. According to the company, fewer than one percent of sites choose to block search bots, while 17 percent enable some method intended to prevent training. Search commonly directs people back to publishers; model training can use their material without producing the same visitor relationship. That difference matters particularly to sites financed through advertising, subscriptions or direct reader engagement.

Cloudflare is changing the definitions of its Bot Management and AI Crawl Control options alongside the launch. Selecting Block for a mixed-use crawler will now stop it altogether, including its search activity. The more limited Disallow AI Training selection permits search crawling while expressing the training restriction. Existing granular configurations that blocked training, or blocked it on pages with advertising, are to be migrated to the new disallow choice where appropriate.

The company says Applebot, Bingbot and Googlebot qualify under the Accountable designation. It also lists relevant Amazon, Anthropic, Meta and OpenAI crawlers as Accountable because those companies separate search and training bots, allowing the latter to be blocked without affecting indexing. Cloudflare notes that Apple already exposes an Applebot-Extended mechanism for declining training.

The current release does not settle every publisher concern. Cloudflare says AI-generated search summaries require more precise controls because the amount of copied or summarized material matters, not merely whether a summary exists. It aims to offer a centralized way for owners to control that use by early next year. Agent crawlers also remain outside the new disallow option while preferences and standards continue to develop.

The result is a more granular policy layer, but compliance still depends on crawler operators honoring declared purposes or being correctly identified and blocked. Cloudflare's announcement describes the controls and commitments; it does not establish that every automated collector on the internet will respect them.