Signal Scanner · ARTIFICIAL INTELLIGENCE & AUTOMATION · 22 August 2026

Blocked by Default: Machine Access to the Web Now Has a Price and a Deadline

With the EU's binding rules aimed at output disclosure and the UK choosing to watch rather than legislate, the terms on which machines may read the web are being left to the market: a default block from 15 September 2026, and pricing moving from per-crawl to per-answer. Publishers first, then anyone whose website feeds somebody else's automation.

Two governments have now said what they intend to do about machine reading of the web, and in both cases it is not this. The EU's binding instrument this cycle governs disclosure of AI output (European Commission, 31/07/2026), and the UK has deliberately proposed to watch the licensing market form rather than legislate into it (GOV.UK, 18/03/2026). Those are choices, not vacuums. The consequence is that commercial terms are being left to settle first, and the earliest concrete default lands on 15 September 2026, when a network serving roughly a fifth of the world's websites starts refusing mixed-purpose crawlers on ad-supported pages (TechCrunch, 01/07/2026).

Signal Identification

A regulatory pivot in sequencing rather than in substance. What is observable is narrow: with legislation held back by design, the first operative access rules and prices are appearing as commercial defaults and brokerage terms. Reading those defaults as precedent-setting is this scan's inference; the sources supply the components and several of them argue the right response is scrutiny, not acceptance.

Time horizon: 1-3 years (default block 15 September 2026; per-use pricing normalised 2027-2028)
terms set 1-2 yrs202620272029
Plausibility band: Medium-High
LowMediumHigh
Geographic / Jurisdictional Scope: United States, EU-27 and United Kingdom primary; Canada, Australia, Japan and Brazil as spillover, reached by the same defaults without a domestic policy decision
PrimaryUSEU-27UK
SpilloverCanadaAustraliaJapanBrazil
Sectors exposed:
News and trade publishingE-commerce and marketplace listingsTravel and hospitality inventoryTechnical documentationMarket-data providersSEO and content operationsEnterprise AI procurement

What's Changing

Cloudflare, used by 20% of websites worldwide, will from 15 September block mixed-use crawlers by default on pages that host ads, covering new customers, new sites from existing customers and all existing free customers (TechCrunch, 01/07/2026). The same announcement priced the waste: over 50% of AI crawl traffic re-fetches unchanged pages. In August the company shifted its default to pay-per-use, paying publishers when content is surfaced in an answer rather than when a bot fetches a file (Press Gazette, 17/08/2026).

The blocking predates the deadline and is already lopsided. Between July 2025 and January 2026 the count of sites actively blocking AI crawlers ran at nearly seven times the count blocking traditional search crawlers such as Googlebot (Towards an Agent-First Web, 17/06/2026), which reads that as evidence the web's access architecture needs redesigning rather than as a case for operator discretion. New is the price attached: roughly 15% of rights-holder revenue at ScalePost, an estimated 30% at Cloudflare, against competitors letting publishers keep 100% and charging the AI company instead (Nieman Lab, 27/05/2026), whose authors argue that spread is itself grounds for regulatory scrutiny of the platform operators setting it.

How much is actually flowing through these rails is still small, and one Tier 2 measure cuts against urgency: weekly use of AI chatbots for news rose 3 percentage points, from 7% to 10%, growth the Reuters Institute calls fast rather than explosive (Digital News Report 2026, 16/06/2026). The gap between modest volumes today and durable terms being written now is the whole of the signal. Brookings makes the timing argument directly: deal structures, price precedents, take rates and governance norms being established at present will be difficult to dislodge once normalised (Brookings, 09/06/2026), and concludes that policy should move early rather than wait.

Intermediary take rates in the emerging licensing market

Intermediary cut taken from rights-holder revenue ScalePost 15% Cloudflare 30% (est.) TollBit, Sphere no cut; fee charged to the AI company 0% 20% 40%

Take rates as reported by Nieman Lab. The Cloudflare figure is an estimate, not a disclosed rate.

Disruption Pathway

Stage one, September 2026 to mid-2027, is a sorting exercise. Crawler operators either split search from training and agent retrieval or lose default reach on ad-supported pages; publishers with deals whitelist their counterparties; everyone else discovers what their bot policy actually says. Stage two, across 2027 and 2028, carries the harder problem: paying per use rather than per fetch obliges the AI firm to report which retrieved documents reached which answers, so answer engines must expose retrieval telemetry to commercial counterparties, years ahead of any statute requiring it. Stage three is generalisation, because once metering rails exist the next customers are not newsrooms but anyone holding proprietary listings, prices or documentation.

Stress concentrates in three places. Firms wanting search discovery without model ingestion cannot reliably separate the two, since Google's crawler serves both purposes, so blocking costs visibility. The long tail has no deal and no bargaining position, converting a default block into lost reach with no offsetting revenue, though Cloudflare argues some local titles are on a path to more licensing revenue than ad revenue (Press Gazette, 17/08/2026). Buy-side enterprises face the third: retrieval costs their AI vendors absorb today become a contracted line item. Crawler policy then becomes a commercial setting owned alongside pricing, and machine-readable content gets valued as a licensable input.

Why This Matters Now

Boards, chief revenue officers and CFOs have three weeks before a default they did not choose changes how their websites behave toward automated readers. Two decisions need an owner and in most organisations neither has one. Who sets crawler policy, now that a single setting governs search visibility, training exposure and a potential revenue line at once? And is the firm's content a marketing asset or a licensable input, given that those answers produce opposite configurations?

Decision-action posture for this signal: Decide — the default changes on a fixed date inside this planning cycle and applies without any action by the site owner, so declining to decide is itself a decision.

Counter-Argument

The strongest objection is that scarcity is being priced before the demand exists. Of those chatbot news users, 42% say they always or often click through to original sources, between social media at 36% and search at 44% (Digital News Report 2026, 16/06/2026). If answer engines are not yet displacing the referral economy, engineered scarcity is a bargaining posture that thins as models route around blocked sites. The UK reached a compatible conclusion on the evidence available to it, proposing not to intervene at this stage (GOV.UK, 18/03/2026).

The objection measures the wrong quantity. Volumes today do not determine who writes the terms, and the terms are being written now: deal structures, price precedents, take rates and governance norms being established at present will be difficult to dislodge once normalised (Brookings, 09/06/2026). Brookings draws the opposite policy conclusion to the UK, arguing calcification is exactly why intervention should come early; this scan takes no position on that, only that the defaults land in September either way.

Implications

Treat this as durable rather than transient, because it changes a default and not a price: prices move back, defaults rarely do. The window runs from September 2026 to mid-2028, closing as per-use accounting settles. Winners are content owners with scarce, verifiable, frequently-updated material, and the intermediaries metering it; losers are firms whose content is substitutable and whose reach depends on being freely readable. Whether the terms now forming should be left to form is contested among the sources: the UK says watch GOV.UK, Brookings and the Open Markets authors say scrutinise now Nieman Lab. This scan takes no side, and notes only that the commercial defaults arrive before that argument is settled.

Early Indicators to Monitor

Disconfirming Signals

Strategic Questions

Keywords

AI crawlers; pay-per-crawl; pay-per-use; machine-readable web; content licensing; crawler blocking; robots.txt; AI Act transparency; retrieval-augmented generation; publisher economics; agentic web; take rates

Bibliography

Source tiers: Tier 1, governments, regulators and intergovernmental bodies. Tier 2, think-tanks, academic institutes, major consultancies and quality data providers. Tier 3, quality journalism and specialist trade press. Tier 4, vendor, company and practitioner sources, used only as directional corroboration.


Prepared by Shaping Tomorrow: 22 August 2026