AI Training on Copyrighted Books: Legal Ambiguity and High Stakes for Tech Companies
August 23, 2026
The legality of training AI models on copyrighted books remains unsettled and will hinge on evolving court interpretations of copyright law and fair use.
Fair use factors—purpose, nature, amount used, and market impact—guide outcomes, with transformative use and lack of direct competition often helping defendants.
Courts have centered fair use in these disputes, and while rulings tend to favor developers when training is framed as familiarization rather than copying, provenance matters.
Anthropic’s settlement of about $1.5 billion, paying roughly $3,000 per work for around 500,000 titles and destroying pirated copies, underscores the high cost of data provenance.
Publishers have filed a separate class action alleging Google overstepped permissions for Gemini training and hid sources by removing copyright management information.
Cases like Thaler v. Perlmutter highlight debates over when AI-generated works are eligible for copyright and how to prove AI involvement.
Thaler v. Perlmutter also established that AI-generated works cannot be copyrighted, leaving many questions about protection for AI-produced content.
Provenance remains a central theme: even winning on fair use can be offset by liability and damages tied to how data was acquired.
The broader litigation landscape shows ongoing risk and significant financial exposure for AI labs as publishers pursue enforcement.
Authors fear livelihood threats from AI tools trained on their works without consent, while some rulings offer cautious reassurance that training can be permissible under certain conditions.
Thomson Reuters v. Ross Intelligence illustrates that training content which directly competes with the original work is unlikely to be fair use.
Google’s case hinges on permission for internal AI training and whether publishers’ materials were used beyond intended limits.
Summary based on 7 sources
Get a daily email with more Tech stories
Sources

TechCrunch • Aug 23, 2026
Is it legal to train AI models on copyrighted books? It’s complicated
CryptoRank • Aug 23, 2026
Is It Legal to Train AI on Copyrighted Books? The Answer Is Complicated
Yahoo • Aug 23, 2026
Is it legal to train AI models on copyrighted books? It’s complicated
Межа • Aug 23, 2026
AI Training on Copyrighted Books Faces Unsettled Legal Questions