
Image credit: Search Engine Journal
NEW YORK – Three publishers and author Scott Turow filed a class action lawsuit against Google on July 10 in New York, alleging the company illegally used copyrighted works to train its Gemini artificial intelligence model.
The lawsuit claims Google copied millions of books and journal articles without authorization from various sources, including Google Books, Play Books, Scholar, and web scrapes, to develop its AI technology.
Hachette Book Group, Cengage Learning, and Elsevier joined novelist Scott Turow and his company S.C.R.I.B.E. as plaintiffs in the U.S. District Court for the Southern District of New York.
The complaint brings four counts against Google, including three for unauthorized reproduction under the Copyright Act and one for removing copyright management information in violation of the Digital Millennium Copyright Act.
Plaintiffs are seeking unspecified damages, an injunction to stop further alleged infringement, an account of all works used, and court orders to delete unauthorized copies of their content.
Internal Google documents cited in the filing reportedly highlighted the potential legal risks of using books from Google Play Books for AI training, with possible fines reaching into the “$10Bs-$100Bs.”
The lawsuit alleges Google acquired content through direct agreements and web scraping from various domains, including pirate sites and paywalled libraries, bypassing traditional crawler controls.
Google’s ‘Google-Extended’ robots.txt token and other established crawler controls do not apply to the specific methods of content acquisition detailed in the complaint, according to the plaintiffs.
The Association of American Publishers, Digital Content Next, Common Crawl Foundation, Anthropic, and Meta were among other organizations mentioned in the context of AI training and content acquisition.
Google published a policy paper on June 25, arguing that training AI models on publicly available web data constitutes a “transformative, non-expressive use” protected under fair use doctrines.
This legal challenge follows increasing scrutiny and a wave of similar lawsuits against AI developers over the use of copyrighted material to train their large language models.
Source: Search Engine Journal
Written by
Joyce de Castro
Joyce is a core team member at Rabbit Rank and the lead author covering SEO news, algorithm updates, industry trends, and actionable ranking strategies.
Keep reading
Related Articles

Google develops ‘Persistent Autonomous Search’ to redefine information retrieval
Google is developing ‘Persistent Autonomous Search,’ a system that keeps queries active until answers are foun...

Retailers unprepared for agentic commerce shift, AI platforms
Agentic commerce is transforming e-commerce, allowing AI-driven transactions. Most top retailers are unprepare...

Continuous Monitoring Key for Brand Visibility in AI Overviews
Continuous monitoring is essential for AI Overviews visibility, surpassing unreliable spot-checks. Strategies...