Cloudflare to Block Mixed-Use AI Crawlers by Default
licensing.
TAGS: Cloudflare, Artificial Intelligence, Web Crawlers, Digital Publishing, AI Training
CONTEUDO:
Cloudflare is set to implement a strict new policy aimed at curbing the unauthorized use of publisher content by artificial intelligence models. Starting September 15, 2026, the company will default to blocking “mixed-use” crawlers—bots that simultaneously perform traditional search indexing and AI model training—from accessing pages containing advertisements.

This update forces a clear separation between tools used for standard search discovery and those used for scraping data to power AI agents or training sets. The new default configuration will apply to all new Cloudflare customers, any site newly onboarded to the platform, and existing users on free plans, unless they manually adjust their security settings to allow such traffic.
Addressing the “Non-Human” Traffic Surge
Matthew Prince, co-founder and CEO of Cloudflare, emphasized the urgency of the move, noting that non-human traffic has recently eclipsed human activity online—a milestone reached earlier than industry projections anticipated. “We must go further and act faster so that a sustainable ecosystem can emerge,” Prince stated regarding the shift.
Cloudflare’s initiative is designed to provide website owners with greater control over their intellectual property. The company argues that while publishers generally support search engine visibility, they are increasingly seeking mechanisms to prevent their content from being harvested for AI training without compensation.
The Push for “Pay Per Use”
Beyond blocking unwanted scrapers, Cloudflare is expanding its monetization tools for publishers. The company is evolving its previous “Pay Per Crawl” marketplace into a broader “Pay Per Use” model. This system allows creators to generate revenue when their content provides tangible value to AI services, rather than simply charging for the initial fetch.
The company is currently testing this framework through partnerships with:
- Ceramic.ai: Publishers receive payment when their content is utilized in AI-generated search results.
- You.com: A system where providers can charge for accessing premium content via the AI platform.
Industry Impact and Google’s Stance
Cloudflare’s data indicates that more than 50% of AI crawl traffic is redundant, spent on re-fetching pages that have not undergone changes. By limiting these bots, the company aims to help publishers preserve bandwidth and compute resources.
The policy specifically challenges the current dominance of major search engines, which Cloudflare suggests maintain an advantage by bundling AI training with standard search discovery. While Google maintains that site owners can opt out of AI training via Google Extended without losing search presence, its primary Googlebot continues to crawl for both search features and AI-integrated responses, a practice Cloudflare’s new tools aim to disrupt.