AI Training on Copyrighted Books Sparks Legal Debate Over Fair Use and Data Provenance Costs
August 23, 2026
The legality of training AI models on copyrighted books remains unsettled, hinging on evolving court interpretations of copyright law and fair use.
Fair use analysis centers on purpose, nature, amount used, and market impact, with transformative use and lack of direct competition shaping outcomes.
Courts have frequently considered fair use in training contexts, with rulings tending to favor developers when training is framed as familiarization rather than copying.
Some observers invoke dataism to describe a society where information becomes the central asset as AI processes and generates knowledge continuously.
Public AI infrastructure and a broad information ecosystem are foregrounded, emphasizing researchers, libraries, archives, and public institutions in knowledge creation and calling for redistributive funding to sustain knowledge infrastructure.
Large Language Models function through data collection, web scraping, tokenization, embeddings, training, and inference, with mechanisms like Retrieval-Augmented Generation supporting outputs.
A settlement of about $1.5 billion in the Anthropic case underscores the high costs of data provenance, including payments around $3,000 per work for hundreds of thousands of works and destruction of pirated copies.
The ANI case is seen as signaling a broader conversation about how information is learned, stored, and accessed in the AI era, beyond any single verdict.
Generative AI learns from vast public data and creates outputs through statistical relationships, challenging traditional ideas of information and expression in copyright law.
There is a shift from information retrieval to computational knowledge creation, necessitating clearer legal definitions of learning, reproduction, transformation, and infringement in AI.
Cases like Thaler v. Perlmutter raise questions about AI-generated works and how to prove AI involvement, affecting copyright ownership and eligibility.
Thaler v. Perlmutter also established that AI-generated works cannot be copyrighted, leaving many questions about protection for AI-produced content unresolved.
Summary based on 8 sources
Get a daily email with more Tech stories
Sources

TechCrunch • Aug 23, 2026
Is it legal to train AI models on copyrighted books? It’s complicated
CryptoRank • Aug 23, 2026
Is It Legal to Train AI on Copyrighted Books? The Answer Is Complicated
Yahoo • Aug 23, 2026
Is it legal to train AI models on copyrighted books? It’s complicated
Межа • Aug 23, 2026
AI Training on Copyrighted Books Faces Unsettled Legal Questions