Photo via Unsplash
Cloudflare gave AI companies a deadline of 15 September 2026. By that date, any company running a crawler has to separate the bot it uses for search indexing from the bot it uses to collect AI training data and to feed AI agents, and declare which is which. Crawlers that have not done so get blocked by default from any page carrying advertising. That is a week from now, and it applies across a network that carries more than a fifth of global web traffic.
The mechanism underneath it is simpler than the policy language. A site behind Cloudflare gets three options for each AI crawler rather than the two the web has offered since the 1990s. Allow it free. Block it. Or charge it a fee per fetch, collected at the network edge before the bot ever loads the page. Cloudflare processes the payment and passes the money on. That third option is pay-per-crawl, and it launched in July 2025 alongside the decision to block AI crawlers by default.
The ratio that made this inevitable
The old bargain of the open web was never written down but everyone understood it. A search engine crawls your pages, sends readers back, and both sides come out ahead. Cloudflare published the numbers on what that exchange had actually become. Google, which has the longest-standing traffic relationship with publishers, crawled about 14 pages for every click it sent back. OpenAI's crawler fetched roughly 1,700 pages per referral. Anthropic's fetched roughly 73,000.
Seventy-three thousand to one is not a degraded bargain. It is a different transaction wearing the old one's clothes. And the volume moved fast enough that it could not be ignored: AI-related crawlers were 22 percent of total crawler requests in spring 2025 and 52 percent by June 2026. Every large language model on the market learned from content collected on those terms. A handful of large publishers sued. Most did not have the resources for a court fight, and the value transferred from the people who produced content to the people who produced models without ever appearing on anyone's balance sheet.
What the arithmetic actually returns
I have watched this story get told to Caribbean boards as a new revenue line, and the arithmetic does not support that for almost anyone in this room. Take a site with a million monthly page views where AI crawlers are one to two percent of traffic, which is typical. At a tenth of a cent per page, that is about 20 dollars a month. At a full cent per page, about 200 dollars. Lunch, not payroll.
There is a second problem, and Cloudflare has conceded it. Pay-per-crawl pays for the fetch, not for the value. A publisher gets paid when a bot retrieves a page whether or not that page ever shaped an answer, and gets nothing when one crawled page becomes the basis of ten thousand responses. The company is developing the feature into Pay Per Use, which would charge when content creates value rather than when it is collected. That is the version worth negotiating against. The current version is measuring the wrong thing accurately.
A control point now exists where none did, and it is worth more than the cents flowing through it. The early adopters of the marketplace, Conde Nast, TIME, the Associated Press, Adweek and Fortune, are not there for the per-crawl money. They are there to establish that access is a negotiated thing.
The Numbers Behind the Meter
- 15 September 2026Deadline for AI companies to separate search crawlers from AI training and agent crawlers, or be blocked by default from ad-bearing pages
- >20%Share of global web traffic sitting behind Cloudflare
- 52%Share of all crawler requests coming from AI crawlers in June 2026, up from 22% in spring 2025
- 73,000 : 1Anthropic crawler fetches per referral sent back; OpenAI roughly 1,700 : 1; Google roughly 14 : 1
- $20 to $200Typical monthly pay-per-crawl revenue on a site with one million monthly page views, at a tenth of a cent and a cent per page
- July 2025Date Cloudflare began blocking AI crawlers by default and launched pay-per-crawl
What the Caribbean actually holds
The regional opportunity here is not publishing. A chamber of commerce metering its newsletter archive would earn about forty dollars a quarter. The opportunity is in long data series that cannot be reconstructed from anywhere else, held by institutions that have never thought of themselves as data owners.
The Caribbean Institute for Meteorology and Hydrology holds regional weather and water records going back decades. National disaster agencies and CDEMA hold post-event damage assessments at a level of granularity that satellite reconstruction cannot match, because somebody walked the site. Agriculture ministries hold yield records across crops and growing conditions that exist nowhere in the global training corpus. Tourism authorities hold arrival and origin patterns going back to the 1990s. Regional insurers hold hurricane loss histories that are, in effect, the only quantitative record of what Category 4 and 5 events do to this specific building stock.
What makes those valuable is not that they are big. By the standards of a training run they are tiny. It is that a model trained overwhelmingly on North Atlantic and continental data performs measurably worse on small island conditions, and this region holds the correction. My own doctoral work on climate modelling for small island states ran into that wall repeatedly: the physics is fine, the regional observations to constrain it are what is missing, and they are sitting in filing systems in Bridgetown and Kingston and St Augustine.
None of these organizations built those records to sell them. That is exactly why the inventory has to happen before the negotiation does.
Two decisions, and neither of them is about price
The first is whether you know what you hold. Almost no board I sit with can name its own data sets, say how far back each one runs, say whether it exists anywhere else, or say who inside the organization controls access. Four questions, written down, per data set. That inventory is the entire prerequisite, and no licensing conversation is possible without it. It is also useful for a dozen reasons that have nothing to do with AI.
The second is whether licensing is even the right answer. For a tourism authority or a private carrier, probably yes. For disaster and public health data, I think the argument often runs the other way. If a model that emergency managers and insurers will use anyway performs badly on Caribbean conditions because Caribbean observations were never in it, the region pays for that in forecasts and premiums, not in licensing fees. Deliberate open publication under clear terms may protect more people than a paywall does. What matters is that it becomes a decision rather than a default, because right now it is being settled by inaction and by whichever crawler got there first.
Who negotiates
Prices in a new market are set by whoever transacts first, and they harden fast. Firms and institutions that reach an agreement in the next year will define what regional data is worth for the decade after. Those that wait will accept whatever rate the market has already settled on, and by then the reference points will have been set by people with no exposure to this region at all.
That is a first-mover advantage with an expiry date on it, which is a phrase I use carefully because most claimed first-mover advantages do not have one. This one does, and the September deadline is the visible edge of it.
Start with the inventory. Four questions per data set, this quarter, before anyone talks to a vendor about a price.
"The meter is not the story. The control point is. For twenty years Caribbean data left the region for free because nobody had a switch. Now there is a switch, and somebody in your organization has to be responsible for it." - Adrian Dunkley, AI Boss
Frequently Asked Questions
What is Cloudflare pay-per-crawl?
Pay-per-crawl is a Cloudflare feature that gives a website owner three choices for each AI crawler rather than two: allow it free, charge a fee each time it fetches a page, or block it entirely. Cloudflare handles the payment processing at the network edge and distributes the revenue, so the decision happens before the bot loads the content. It launched in July 2025 alongside Cloudflare's move to block AI crawlers by default.
What changes on 15 September 2026?
Cloudflare gave AI companies until 15 September 2026 to separate the crawlers they use for search from the crawlers they use for AI training and AI agents, and to declare which purpose each crawler serves. After that date, Cloudflare's default settings block mixed-use crawlers from any page that carries advertising. The practical effect is that a crawler which will not say what it is collecting content for loses default access to the commercial web.
How much AI crawler traffic is there?
AI-related crawlers accounted for 52 percent of all crawler requests in June 2026, up from 22 percent in spring 2025. Cloudflare's own data on the exchange behind that traffic is what made the policy case: OpenAI's crawler fetched roughly 1,700 pages for every single referral it sent back to a site, Anthropic's fetched roughly 73,000 per referral, and Google, which has the most established traffic relationship with publishers, crawled about 14 times for every click it returned.
How much money can a publisher actually make from pay-per-crawl?
For most publishers, very little. On a site with a million monthly page views where AI crawlers are one to two percent of traffic, a price of a tenth of a cent per page produces roughly 20 dollars a month and a price of one cent per page produces roughly 200 dollars. Neither figure funds anything. The publishers for whom the numbers work are large outlets with heavy crawler traffic and content that models specifically want, and the early adopters of the marketplace reflect that: Conde Nast, TIME, the Associated Press, Adweek and Fortune.
What is the main criticism of pay-per-crawl?
It pays for the fetch rather than for the value. A publisher is compensated when a bot retrieves a page whether or not that page ever influenced a model's output, and gets nothing when a single crawled page becomes the basis of thousands of answers. Cloudflare has acknowledged the gap and is developing the feature into Pay Per Use, which would charge AI companies when content creates value rather than when it is collected. Until that ships, the metering measures the wrong thing.
Which Caribbean data would AI companies actually pay for?
Long time series that cannot be reconstructed from anywhere else and that global models are demonstrably thin on. Regional meteorological and hydrological records held by institutions such as the Caribbean Institute for Meteorology and Hydrology, post-disaster damage assessments held by national agencies and CDEMA, agricultural yield records held by ministries, visitor arrival and origin patterns held by tourism authorities, and insurance loss histories held by regional carriers. What makes these valuable is not volume. It is that a model trained mostly on North Atlantic and Pacific data performs measurably worse on small island conditions, and nobody else holds the correction.
Should a Caribbean organization license its data or not?
Ask two questions before anything else. Is the data unique, meaning it cannot be reconstructed from public sources, and is it a series rather than a snapshot? And is there a public-interest reason it should stay open, which is often true of disaster and health data where wider model accuracy directly protects people. If the answer to the first is no, there is no asset to license. If the answer to the second is yes, the right move may be deliberate open publication under clear terms rather than a price, so that the regional correction reaches the models that need it.
What should a Caribbean board do about this before the end of the quarter?
Find out what your organization actually holds, in a written inventory naming each data set, its time span, whether it exists anywhere else, and who inside the organization controls it. Most boards cannot answer those four questions today, and no licensing conversation is possible without them. Then decide, deliberately, whether each set is open, licensed, or closed. The decision matters more than the price, and it is a decision that is currently being made by default and by inaction.