The AI training-data fight has reached Google Books. Hachette Book Group, Cengage Learning, Elsevier and novelist Scott Turow have filed a class action against Google in the U.S. District Court for the Southern District of New York, alleging Gemini was trained on books obtained through Google Books in violation of an agreement that permitted only snippet display. The complaint cites an internal Google document indicating the company faced “$10Bs-$100Bs” in potential fines — offered by plaintiffs as evidence Google understood the legal risk. These allegations are unproven, and Google has not been found liable of anything.
The numbers in this piece
01The claim, and the number hanging over it
The alleged breach is specific rather than general: not that scraping books is inherently unlawful, but that Google had a negotiated arrangement limiting use to snippets and then used the same corpus to train a commercial model. That framing matters. Contract-adjacent claims are often harder for a defendant to wave away with a broad fair-use argument than pure copyright claims are.
The plaintiff group is close to the one that filed a similar action against Meta in May 2026. And there is already a benchmark on the board: Anthropic previously settled a comparable piracy class action with authors for $1.5 billion. That figure is now the anchor every plaintiff’s lawyer in this space starts from — and the reason the internal “$10Bs-$100Bs” line, if it holds up in discovery, is the most consequential detail in the filing.
02Why Google is the outlier here
Google has notably refused to strike licensing deals with digital publishers, unlike several of its competitors. That refusal is turning into a strategic problem. Some publishers — USA Today among them — are reportedly considering delisting from Google Search within six to 12 months, which until very recently would have read as self-harm rather than strategy.
That such a move is even discussable tells you how much the traffic calculus has changed. When search referrals decline enough, the cost of leaving falls toward the value of staying, and litigation stops being the only lever publishers have.
03Why this matters
| Contract terms may prove sharper than copyright claims | If this case advances on the theory that Google exceeded a negotiated permission, then every agreement you have with a platform — archive access, snippet rights, syndication, API terms — is a potential constraint on AI training. Go read them. |
|---|---|
| The settlement benchmark is now public | $1.5 billion from Anthropic reframes what licensing negotiations are worth. Even if you'll never litigate, the number sets the floor under what your corpus should command in a deal. |
| Google's no-licensing posture is a durable risk, not a phase | Competitors are paying. Google is not, and is being sued instead. Plan your Google-dependent revenue on the assumption that no licensing check is coming. |
A commercial model built on a corpus obtained under narrower terms is the cleanest version of the argument publishers have been making for three years.
04What publishers should do
05What marketers should do
06The bottom line
A commercial model built on a corpus obtained under narrower terms is the cleanest version of the argument publishers have been making for three years. Whether it wins is genuinely uncertain — these are allegations, not findings. But the combination of a $1.5 billion settlement benchmark, an internal document about tens of billions in exposure, and publishers seriously weighing life without Google Search means the negotiating table looks different than it did last quarter. Know what your rights actually say before someone else decides for you.