Publishers, author sue Google over AI training with copyrighted books

Joyce de Castro Joyce de Castro · · 2 min read

Share this article

NEW YORK – Three publishers and author Scott Turow filed a class action lawsuit against Google on July 10 in New York, alleging the company illegally used copyrighted works to train its Gemini artificial intelligence model.

The lawsuit claims Google copied millions of books and journal articles without authorization from various sources, including Google Books, Play Books, Scholar, and web scrapes, to develop its AI technology.

Hachette Book Group, Cengage Learning, and Elsevier joined novelist Scott Turow and his company S.C.R.I.B.E. as plaintiffs in the U.S. District Court for the Southern District of New York.

The complaint brings four counts against Google, including three for unauthorized reproduction under the Copyright Act and one for removing copyright management information in violation of the Digital Millennium Copyright Act.

Plaintiffs are seeking unspecified damages, an injunction to stop further alleged infringement, an account of all works used, and court orders to delete unauthorized copies of their content.

Internal Google documents cited in the filing reportedly highlighted the potential legal risks of using books from Google Play Books for AI training, with possible fines reaching into the “$10Bs-$100Bs.”

The lawsuit alleges Google acquired content through direct agreements and web scraping from various domains, including pirate sites and paywalled libraries, bypassing traditional crawler controls.

Google’s ‘Google-Extended’ robots.txt token and other established crawler controls do not apply to the specific methods of content acquisition detailed in the complaint, according to the plaintiffs.

The Association of American Publishers, Digital Content Next, Common Crawl Foundation, Anthropic, and Meta were among other organizations mentioned in the context of AI training and content acquisition.

Google published a policy paper on June 25, arguing that training AI models on publicly available web data constitutes a “transformative, non-expressive use” protected under fair use doctrines.

This legal challenge follows increasing scrutiny and a wave of similar lawsuits against AI developers over the use of copyrighted material to train their large language models.


Joyce de Castro

Written by

Joyce de Castro

Joyce is a core team member at Rabbit Rank and the lead author covering SEO news, algorithm updates, industry trends, and actionable ranking strategies.

Keep reading

Related Articles

Ready to Dominate Search Results?

Let our experts analyze your website and create a custom SEO strategy that drives real results.