AI & the open web

European Publishers Take Four Times the AI Scraping — and Robots.txt Is Ignored

APH Publisher Desk ·5 min read Share Print
In this piece
    Figure Meanwhile the compensating traffic simply is not arriving
    0.05%of European publishers' external referrals came
    Publisher Desk

    European publishers have spent a year describing the AI scraping problem anecdotally. This week, TollBit put numbers on it — and the numbers are worse than the anecdotes.

    TollBit analysed data from 40 scraping vendors across 3,906 publishers, 456 of them European, and found a consistent regional imbalance. Median AI scrapes per site were four times higher on European sites than on North American ones. European publishers received one human visit for every 179 bot visits. And their robots.txt no-scrape instructions were ignored nearly three times more often than those of North American sites — meaning the region taking the most scraping also has the least effective defences against it.

    The direction of travel is as troubling as the level. Europe’s scrape-to-referral ratio is now 227:1, against 150:1 in Q1 — a 50% worsening in a single quarter. Meanwhile the compensating traffic simply is not arriving: just 0.05% of European publishers’ external referrals came from AI apps, versus 0.16% in North America. Digiday reported the findings on 14 August.

    The old bargain of the open web — we crawl you, we send you traffic — was always thin for AI crawlers. At 227:1 and falling, in Europe it is now barely detectable.

    01A claim that directs your measurement, not one that settles it

    Before acting on these figures, note who produced them. TollBit sells scraping monetisation — its business is converting bot traffic into paid access, which means it benefits when the scraping problem looks large and unmanaged. That does not make the data wrong, but it makes it a vendor claim, and other infrastructure providers with visibility into the same traffic see different pictures: Cloudflare and DataDome have reported different regional patterns. Methodology matters enormously here — which sites are in the panel, how bots are classified, what counts as a scrape versus a fetch — and none of the competing datasets is neutral.

    The correct operator response is neither to dismiss the report nor to repeat it as fact. It is to treat it as a hypothesis your own logs can test. Every publisher already possesses the ground truth about its own scraping: server logs, CDN reports, bot-management telemetry. The regional headline is interesting; your number is actionable.

    The robots.txt finding deserves particular attention, because it exposes the gap between policy and enforcement. A no-scrape directive that is ignored three times as often is not a defence — it is a request. If your robots.txt and your server logs disagree about which crawlers are being served content, you have a fixable engineering problem, and fixing it is the difference between having a crawler policy and having a wish. This is exactly the gap the Stealth Bot Prohibition Act — covered elsewhere in this issue — is designed to close legislatively, but no publisher should wait for legislation to close it operationally.

    02The economics of asymmetric scraping

    For European publishers specifically, three consequences flow from a documented 4x imbalance. First, cost: bots consume bandwidth, CDN egress and compute whether or not they ever send a reader back, and at 179 bot visits per human visit, the infrastructure bill for serving machines can materially exceed the bill for serving your actual audience. Second, leverage: EU enforcement mechanisms — from the AI Act’s transparency provisions to copyright frameworks — run on evidence from affected parties, and a publisher who can document the imbalance is worth more to a trade-body submission than ten who can describe it. Third, negotiation: a scraper who wants licensed access negotiates differently against a publisher who can show precisely what unlicensed access has been costing.

    03Why this matters for publishers

    The traffic bargain is dead in the data, not just the discourseAt 227:1 scrape-to-referral and 0.05% of referrals from AI apps, there is no longer a serious argument that European publishers are being compensated in traffic for what crawlers extract. That reframes every build-or-block decision.
    Robots.txt non-compliance is now quantifiedNearly three times the ignore rate means the polite-request era of crawler governance is measurably over in Europe. Enforcement has to move to the server level — blocking, throttling, metering — where compliance is not optional.
    Bot load is contaminating your audience dataAt these ratios, any analytics or audience figure that has not rigorously separated bot from human traffic is wrong, and buyers will eventually ask how you separate them. Better to have the answer before the question.
    Vendor-supplied numbers cut both waysTollBit's figures strengthen the publisher case, but building your licensing posture on a monetisation vendor's panel data is a fragile foundation. Your own logs are unimpeachable; use theirs for context, yours for negotiation.

    04What publishers should do

    05The bottom line

    The value of this report is not its precision — vendor panels are never precise — but its direction, which no competing dataset disputes: scraping is growing, referrals are not, and the imbalance is worst where enforcement is weakest. For European publishers the strategic conclusion is uncomfortable and clarifying at once. The traffic-for-content trade is no longer a trade you need to preserve, which means blocking, metering and licensing are no longer risks to discovery — they are the only rational responses to a counterparty that has stopped paying in the only currency it ever offered. Measure your own logs, cost your own scraping, and negotiate from your own evidence.

    Sources & caveats

    Sources: Digiday, “European publishers are getting hit harder by AI bot scraping, report finds” (14 August 2026), reporting TollBit data covering 40 scraping vendors and 3,906 publishers, 456 of them European. All headline figures — the 4x scrape multiple, 179:1 bot-to-human visits, ~3x robots.txt ignore rate, 150:1 to 227:1 ratio deterioration, and 0.05% vs 0.16% AI-app referral shares — are TollBit’s, and TollBit sells scraping monetisation, so the figures are vendor-supplied and methodology-dependent; Cloudflare and DataDome have reported different regional patterns from their own vantage points.

    The weekly

    One letter a week, from the desk that runs the auctions.

    What actually moved in yield, CTV and curation across our publishers — written by the people who saw it, not a content team. No digests, no roundups, one email.

    One email a week. Unsubscribe in one click. We never share or sell the list.

    More from this issue

    Ran alongside this piece in the Weekly of 16 August 2026 — read the whole issue →