By Raluca Mitchell.
Second in a series from the Prebid LLM Taskforce.
Somewhere around mid-September of this year, Cloudflare changed the default for roughly a fifth of the internet: ad-supported sites on its network now block AI training and agent crawlers automatically while still admitting search crawlers, unless the owner sets otherwise. Two years ago the only widely available lever was a text file that stated a preference; today the front door can tell training from search from agents and treat each differently.
A default is a starting point, not a strategy. As the first article in this series showed, 39% to 46% of crawler traffic on Duplex’s properties declares nothing and can’t be classified: not verified enough to bill, not identifiable enough to block without also catching real readers. That tier is why “block all” and “allow all” are both incomplete answers. The question a publisher faces isn’t which way to flip a switch. It’s what it will accept, from whom, in exchange for what, and how to enforce that with the mechanisms available today.
robots.txt declares; the edge enforces
A July 2026 audit of 10,894 domains by HasData found that 39.5% of sites disallowing GPTBot in robots.txt still served it a live page. Where blocks were enforced, they rested on one header: from the same IP, publisher sites served a plain browser 83.8% of the time and GPTBot 54.2% of the time. A crawler doesn’t need to rotate an IP to get through. It only needs to stop calling itself GPTBot.
robots.txt, in other words, is a declaration with no enforcement behind it. It binds only the actors who choose to honor it. Enforcement happens at the edge, in the CDN, WAF, or bot-management layer, where a request can actually be refused.
Identity is the harder half. In late August, GreyNoise found a scanning campaign from 824 IP addresses forging the names of 13 AI crawlers while probing for exposed credentials. A user-agent string alone establishes little. What publishers can do now is allowlist known partners on user agent, published IP ranges, and network ASN together, and challenge anything presenting a known name from an unexpected network. Publishers running such allowlists report catching “Googlebot” traffic arriving from networks Google doesn’t operate. The method needs maintenance, but it runs on tools most publishers already have.
The longer-term approach is cryptographic. Under Web Bot Auth, a proposed IETF standard, a crawler signs each request with a private key and publishes the public half at a known address, so the server checks a signature instead of trusting a string. Cloudflare, Akamai, AWS WAF, Vercel and HUMAN already verify these signatures, and OpenAI’s ChatGPT agent signs its requests.
Seven behaviors, one decision each
Identity, even verified, is only half the picture. In one week on qz.com, Bingbot pulled 761,275 pageviews, about 2.9 times Googlebot’s, and that single user agent feeds a search index and an AI answer engine at the same time. One identity, two different economic relationships, and no mechanism to offer different terms to each. Purpose has to be declarable, not only identity.
Seen that way, “AI traffic” is at least seven distinct behaviors, and each publisher decides for its own business whether each one is welcome, tolerated, priced, or refused:
- Indexing a headline and a link
- Retrieving a page at a user’s request, either bringing the reader to the page or reading it on their behalf without a visit
- Quoting or citing content inside an answer, with or without a referral
- Producing a substantial answer from the content
- Training a model on the archive
- Executing a subscription or purchase on a reader’s behalf
- Undeclared automated scraping, including intermediaries repackaging content for AI use without identifying themselves or paying
This article offers no verdict on any of them, because the verdicts differ by business. A subscription publisher may welcome the agent that completes a checkout and decline the one that summarizes; a publisher pursuing visibility in generated answers may want to be read so that it is cited. One useful lens is distribution versus substitution: does the behavior return value to the publisher’s own experience, or replace it? Answered per behavior and written down, those answers are business rules. Enforced at the edge, they are a policy.
The rules follow the business model
Which behaviors count as distribution depends on how a publisher is paid. E-commerce reached this fork first: Amazon has blocked shopping agents, most recently Meta’s Muse; Shopify has opened its merchants’ storefronts to them. Same technology, two different bets on what endures, control of the customer and the interface, or participation in the ecosystem agents run on.
“A publisher could choose either one of those two things,” says Armando Roggio, VP of Audience Development at Quartz Media Network. “Is it better to be part of the ecosystem, or is it more important to control the customer and the interface? The answer depends on how you monetize.”
If revenue comes through licensing and syndication, an agent reading the page is the business functioning as designed. If revenue depends on a person seeing an ad, an agent reading without a human behind it is revenue that didn’t arrive. The same behavior is distribution under one model and substitution under the other, which is why no outside party can write a publisher’s rules for it, and why rules written only around today’s revenue may protect a model that is shrinking.
Where to start
None of this needs new technology. The front door a publisher already pays for, usually a CDN or a bot-management service, can see who is visiting, verify who they are, and treat each kind differently. What it usually lacks is the rules. Five steps, most of them an afternoon’s work with whoever runs the website or the vendor behind it, close that gap:
- Look at what is visiting. Pull a week of logs and separate traffic that identifies itself from traffic that doesn’t.
- Check identity by more than the name. Confirm trusted crawlers by user agent, IP range and network together, and challenge any that don’t match.
- Decide on the seven behaviors. Mark each one welcome, tolerated, priced or refused for your business, and write it down.
- Keep robots.txt and add the edge. Crawlers that honor robots.txt still follow it. For those that don’t, set matching rules at the CDN, where they are enforced.
- Set a default for traffic that won’t identify itself. Challenge it, slow it down, or let it through and watch it, but make it a decision rather than an accident.
One caution before tightening anything: keep the contextual, brand-safety and verification crawlers your ad stack depends on allowlisted, or the bids they support leave with them.
The industry has spent two years asking whether to block AI. A publisher that works through these five steps has a better answer: a written policy its own edge enforces.
In the following articles, we’ll look more closely at the emerging content monetization ecosystem and the options publishers have for turning AI access into revenue. Read the first article in the series for a closer look at what publishers are actually seeing in AI crawler traffic, referrals, and costs.
Sources
Cloudflare’s new AI traffic categories (Search, Agent, Training) and the September 15, 2026 default that blocks Training and Agent crawlers on ad-monetized pages while allowing Search: Jin-Hee Lee and Bryan Becker, “Your site, your rules: new AI traffic options for all customers,” Cloudflare blog, July 1, 2026,
https://blog.cloudflare.com/content-independence-day-ai-options/
Duplex network data on the 39% to 46% unclassified share of crawler traffic and the Bingbot and Googlebot pageview comparison on qz.com: Duplex internal traffic analysis, one week, September 2026, as reported in the first article in this series, “The Current State, from a Publisher’s Point of View.”
HasData audit of 10,894 domains (robots.txt compliance rates and same-IP serving rates): HasData, “AI Crawler Block Index,” July 2026, https://hasdata.com/blog/ai-crawler-block-index
GreyNoise scanning campaign impersonating 13 AI crawlers from 824 IP addresses: GreyNoise, “Threat actors posing as AI crawlers,” August 2026, https://www.greynoise.io/blog/threat-actors-posing-as-ai-crawlers
Web Bot Auth standard and Cloudflare’s Verified Bots program: Cloudflare, “Web Bot Auth,” https://blog.cloudflare.com/web-bot-auth/

Leave a Reply
You must be logged in to post a comment.