AI & the open web

Stealth Crawlers Are Routing Around Your Robots.txt

APH Publisher Desk ·4 min read Share Print
In this piece
    Figure AI scraper traffic grew 597% across 2025
    AI SCRAPER TRAFFIC GROWTHAI scraper traffic grew 597% across 2025. Indexed so the before value is 100; after is 697.+597%AI SCRAPER TRAFFIC GROWTHINDEXED · BEFORE = 100100697BEFOREAFTER
    Publisher Desk

    Every publisher strategy for getting paid by AI companies — blocking, pay-per-crawl, tiered licensing — rests on one assumption: that bots identify themselves. This week Digiday laid out, in plain terms, how large the population of bots that simply don’t has become. Stealth crawlers mask their identity, ignore robots.txt and mimic human browsing patterns via residential IPs. They are, by design, invisible to the tools most publishers rely on — which means your blocklist and your licensing deals are both being routed around.

    The scale numbers are the story. Cloudflare now puts bots at over half of all web traffic. AI scraper traffic grew 597% across 2025. Scraping accounts for nearly 20% of site traffic at the median organisation. And People Inc. alone blocks more than 30,000 user agents a day — a figure that describes both the size of the problem and the size of the operational burden of fighting it.

    The policy response is forming around identification. News/Media Alliance CEO Danielle Coffey framed the fix: “Bad actors, bad bots must identify themselves, and then when they do, we can stop them.” New York’s Stealth Crawler Prohibition Act attaches penalties of up to $15,000 a day per violation, and a federal counterpart — the Stealth Bot Prohibition Act — was introduced in the House just after this news week closed, following New York’s lead. Law moves slower than scrapers, though. For now, enforcement is an engineering problem that lives in your infrastructure budget.

    The numbers in this piece

    597%AI scraper traffic growth
    20%of site traffic at the median
    $15KNew York's Stealth Crawler

    01A request is not a control

    It is worth being precise about what robots.txt actually is: a text file expressing a preference, honored voluntarily by crawlers that choose to read it. Every monetisation strategy built on top of it inherits that voluntariness. A blocklist stops the crawlers polite enough to announce themselves — which increasingly means it stops the crawlers you could have negotiated with, while the ones extracting value anonymously sail through disguised as readers in residential IP space.

    That asymmetry has a nasty commercial consequence. If licensed access can be undercut by unlicensed scraping that carries no cost and no identity, the licensed path is competing against free. Publishers negotiating content deals need to be able to demonstrate that the unlicensed path is expensive, slow and unreliable — otherwise the deal is priced against a leak. Which is why the practical conclusion of the piece matters more than the alarm: latency, whitelisting and monitoring beat blocklists.

    The deeper shift is from a default-open posture to a default-deny one. A blocklist fails open every time a new stealth agent appears, which is daily. An allowlist — explicit permission for the crawlers you have chosen: search, verified partners, licensed AI — fails safe. Combined with rate-limiting and tarpitting of suspect traffic, it changes the economics for the scraper: every disguised request costs them compute and time, without requiring you to win an identification argument first.

    02Why this matters for publishers

    Your blocking strategy may already be theaterIf a meaningful share of AI crawling arrives disguised, the crawlers named in your robots.txt and your bot-management console are not the population taking your content — they are the subset polite enough to be counted.
    Licensing leverage leaks through the same holeA pay-per-crawl or licensing tier is a price on identified access. Anonymous access at zero undercuts it structurally, and sophisticated counterparties on the other side of the table know it.
    Bot management is now a real line itemPeople Inc. is fighting 30,000 user agents a day. That fight has headcount, tooling and CDN costs attached, and those costs belong in any calculation of what a licensing deal is worth to you.
    Your audience numbers are exposed tooIf more than half of web traffic is automated, buyers will eventually ask how much of yours is. A publisher who cannot produce a clean human-only traffic figure will have one produced for them, by a verification vendor, on less friendly terms.
    The stealth crawler story is not really a story about bad bots; it is a story about a load-bearing assumption failing.
    Figure 2 Scraping accounts for nearly 20% of site traffic at the median organisation
    20%of site traffic at the median
    Publisher Desk

    03What publishers should do

    04The bottom line

    The stealth crawler story is not really a story about bad bots; it is a story about a load-bearing assumption failing. The entire emerging market for paid AI access — blocks as leverage, crawls as billable events, licenses as tiers — assumes the counterparty shows ID. A growing share simply doesn’t, and the gap between the bots you govern and the bots you get is where publisher leverage drains away. Legislation may eventually force identification, and Coffey’s framing points the right direction. But no publisher should run a 2026 content business on the timeline of a bill. Default-deny, allowlist what you trust, tax what you can’t identify, and know your own numbers — because every strategy downstream of “bots identify themselves” is only as strong as the percentage that actually do.

    Sources & caveats

    Sources: Digiday, “WTF is a stealth crawler?” (August 4, 2026), including the Danielle Coffey quote, the Cloudflare bot-share figure, the 597% AI scraper growth figure, the ~20% median scraping share, the People Inc. 30,000-user-agents-a-day figure and the New York Stealth Crawler Prohibition Act penalties. Scale figures originate with Cloudflare and other infrastructure vendors and are vendor-measured, not independently audited. The federal Stealth Bot Prohibition Act reference is from AdExchanger (August 10, 2026) — one day outside this news week, included as the federal follow-through on the same story.

    The weekly

    One letter a week, from the desk that runs the auctions.

    What actually moved in yield, CTV and curation across our publishers — written by the people who saw it, not a content team. No digests, no roundups, one email.

    One email a week. Unsubscribe in one click. We never share or sell the list.

    More from this issue

    Ran alongside this piece in the Weekly of 9 August 2026 — read the whole issue →