On July 1, Cloudflare drew a line that much of the AI industry will have to reckon with by autumn. The company, which sits in front of a large share of the world's web traffic, told AI firms they have until September 15 to sort their crawlers into clear categories or risk being blocked from a big slice of the internet. The policy was reported by TechCrunch and laid out on Cloudflare's own blog.
The heart of it is a demand for honesty about purpose. Cloudflare wants site owners to be able to treat three kinds of automated visitor differently. A search crawler indexes a page so it can point people back to it later. An agent acts in real time on behalf of a user, fetching a page to answer a question in the moment. A training crawler gathers content to feed a model. Until now these have often arrived under the same banner, and a publisher had little way to welcome one while turning another away.
From September 15, the defaults change. For new customers, and for new sites added by existing ones, Cloudflare will allow search but block training and agent traffic on any page that carries advertising. The same defaults will apply to existing free customers who have not adjusted their own settings by that date. The company's stated aim is blunt: publishers should be paid, or at least see referral traffic, when their work is used.
The mixed-use problem
The sharpest part of the policy targets what Cloudflare calls mixed-use crawlers. These are bots that scrape for several purposes at once and give site owners no way to separate them. Under the new rules, such crawlers will be blocked outright on pages that run ads. That is a pointed message to the largest players. Apple, Google and Microsoft's Bing all run crawlers that could fall on the wrong side of the line, though each offers an AI opt-out that may let it avoid the block.
For years the bargain of the open web was simple enough. Search engines took your content and sent readers back in return. AI answer engines have strained that bargain, because a chatbot that summarises your article hands you nothing. Cloudflare is trying to write a new contract, one where the type of use is declared up front and priced accordingly.
Whether it holds
It is not a law, and Cloudflare cannot force anyone to reclassify anything. What it can do is make scraping expensive in reach. A crawler that refuses to declare itself loses access to a meaningful portion of the ad-supported web, and for a model that lives on fresh content, that stings. The company is betting that this leverage is enough to change behaviour where polite requests, such as the long-ignored robots.txt file, never could.
The open question is what the AI labs do next. They can play along and label their bots, they can lean on their opt-outs, or they can look for ways around the wall. Which of those they choose, over the next two months, will tell us a good deal about who really holds the whip hand between the companies that make the models and the ones that guard the content those models need.
Sources
- i. techcrunch.com
- ii. blog.cloudflare.com
- iii. www.nbcnews.com
- iv. www.theregister.com
Commentarii · 0