AI Firms Buy Old Books for Training, Face Content Provenance Issues

Joyce de Castro Joyce de Castro · · 2 min read

Share this article

Global artificial intelligence companies are purchasing physical books for training data while also generating and monetizing AI-created content, raising concerns about content provenance and detection.

This approach highlights a paradox where firms seek human-generated data for foundational models but struggle to identify and manage the synthetic material their systems produce.

Companies using services like ISBNdb are acquiring older printed books, valuing their content for training before the widespread proliferation of AI-generated text, according to industry reports.

Traditional search analytics are largely ineffective for discerning AI-generated answers due to their dynamic and synthetic nature, complicating efforts to measure and track their origin.

Google announced its SynthID watermarking technology for 100 billion AI-generated images and videos, along with 60,000 years of audio, partnering with OpenAI, Kakao, ElevenLabs, and NVIDIA.

Anthropic committed to embedding watermarks in all text generated by its Claude AI globally starting August 2, 2026, a move driven by compliance with the European Union AI Act.

However, published watermarking methods have documented weaknesses against techniques like rewriting, translation, and short output generation, researchers said.

Google’s public SynthID text code is currently designated for research purposes, not production use, indicating ongoing development in the field.

The company has historically maintained secrecy around its spam detection mechanisms, suggesting its true AI text detection methods will likely remain undisclosed.

Despite these efforts, watermark removers for AI models including Claude, Gemini, and OpenAI have already appeared on platforms like GitHub, illustrating the persistent challenge of evasion.

This ongoing situation between AI content generation and detection raises fundamental questions about the future of digital provenance and content authenticity.


Joyce de Castro

Written by

Joyce de Castro

Joyce is a core team member at Rabbit Rank and the lead author covering SEO news, algorithm updates, industry trends, and actionable ranking strategies.

Keep reading

Related Articles

Ready to Dominate Search Results?

Let our experts analyze your website and create a custom SEO strategy that drives real results.