European publishers have spent a year describing the AI scraping problem anecdotally. This week, TollBit put numbers on it — and the numbers are worse than the anecdotes.
TollBit analysed data from 40 scraping vendors across 3,906 publishers, 456 of them European, and found a consistent regional imbalance. Median AI scrapes per site were four times higher on European sites than on North American ones. European publishers received one human visit for every 179 bot visits. And their robots.txt no-scrape instructions were ignored nearly three times more often than those of North American sites — meaning the region taking the most scraping also has the least effective defences against it.
The direction of travel is as troubling as the level. Europe’s scrape-to-referral ratio is now 227:1, against 150:1 in Q1 — a 50% worsening in a single quarter. Meanwhile the compensating traffic simply is not arriving: just 0.05% of European publishers’ external referrals came from AI apps, versus 0.16% in North America. Digiday reported the findings on 14 August.
The old bargain of the open web — we crawl you, we send you traffic — was always thin for AI crawlers. At 227:1 and falling, in Europe it is now barely detectable.
01A claim that directs your measurement, not one that settles it
Before acting on these figures, note who produced them. TollBit sells scraping monetisation — its business is converting bot traffic into paid access, which means it benefits when the scraping problem looks large and unmanaged. That does not make the data wrong, but it makes it a vendor claim, and other infrastructure providers with visibility into the same traffic see different pictures: Cloudflare and DataDome have reported different regional patterns. Methodology matters enormously here — which sites are in the panel, how bots are classified, what counts as a scrape versus a fetch — and none of the competing datasets is neutral.
The correct operator response is neither to dismiss the report nor to repeat it as fact. It is to treat it as a hypothesis your own logs can test. Every publisher already possesses the ground truth about its own scraping: server logs, CDN reports, bot-management telemetry. The regional headline is interesting; your number is actionable.
The robots.txt finding deserves particular attention, because it exposes the gap between policy and enforcement. A no-scrape directive that is ignored three times as often is not a defence — it is a request. If your robots.txt and your server logs disagree about which crawlers are being served content, you have a fixable engineering problem, and fixing it is the difference between having a crawler policy and having a wish. This is exactly the gap the Stealth Bot Prohibition Act — covered elsewhere in this issue — is designed to close legislatively, but no publisher should wait for legislation to close it operationally.
02The economics of asymmetric scraping
For European publishers specifically, three consequences flow from a documented 4x imbalance. First, cost: bots consume bandwidth, CDN egress and compute whether or not they ever send a reader back, and at 179 bot visits per human visit, the infrastructure bill for serving machines can materially exceed the bill for serving your actual audience. Second, leverage: EU enforcement mechanisms — from the AI Act’s transparency provisions to copyright frameworks — run on evidence from affected parties, and a publisher who can document the imbalance is worth more to a trade-body submission than ten who can describe it. Third, negotiation: a scraper who wants licensed access negotiates differently against a publisher who can show precisely what unlicensed access has been costing.
03Why this matters for publishers
| The traffic bargain is dead in the data, not just the discourse | At 227:1 scrape-to-referral and 0.05% of referrals from AI apps, there is no longer a serious argument that European publishers are being compensated in traffic for what crawlers extract. That reframes every build-or-block decision. |
|---|---|
| Robots.txt non-compliance is now quantified | Nearly three times the ignore rate means the polite-request era of crawler governance is measurably over in Europe. Enforcement has to move to the server level — blocking, throttling, metering — where compliance is not optional. |
| Bot load is contaminating your audience data | At these ratios, any analytics or audience figure that has not rigorously separated bot from human traffic is wrong, and buyers will eventually ask how you separate them. Better to have the answer before the question. |
| Vendor-supplied numbers cut both ways | TollBit's figures strengthen the publisher case, but building your licensing posture on a monetisation vendor's panel data is a fragile foundation. Your own logs are unimpeachable; use theirs for context, yours for negotiation. |
04What publishers should do
05The bottom line
The value of this report is not its precision — vendor panels are never precise — but its direction, which no competing dataset disputes: scraping is growing, referrals are not, and the imbalance is worst where enforcement is weakest. For European publishers the strategic conclusion is uncomfortable and clarifying at once. The traffic-for-content trade is no longer a trade you need to preserve, which means blocking, metering and licensing are no longer risks to discovery — they are the only rational responses to a counterparty that has stopped paying in the only currency it ever offered. Measure your own logs, cost your own scraping, and negotiate from your own evidence.