AI Training on Copyrighted Books: Legal Ambiguity and High Stakes for Tech Companies

August 23, 2026
AI Training on Copyrighted Books: Legal Ambiguity and High Stakes for Tech Companies
  • The legality of training AI models on copyrighted books remains unsettled and will hinge on evolving court interpretations of copyright law and fair use.

  • Fair use factors—purpose, nature, amount used, and market impact—guide outcomes, with transformative use and lack of direct competition often helping defendants.

  • Courts have centered fair use in these disputes, and while rulings tend to favor developers when training is framed as familiarization rather than copying, provenance matters.

  • Anthropic’s settlement of about $1.5 billion, paying roughly $3,000 per work for around 500,000 titles and destroying pirated copies, underscores the high cost of data provenance.

  • Publishers have filed a separate class action alleging Google overstepped permissions for Gemini training and hid sources by removing copyright management information.

  • Cases like Thaler v. Perlmutter highlight debates over when AI-generated works are eligible for copyright and how to prove AI involvement.

  • Thaler v. Perlmutter also established that AI-generated works cannot be copyrighted, leaving many questions about protection for AI-produced content.

  • Provenance remains a central theme: even winning on fair use can be offset by liability and damages tied to how data was acquired.

  • The broader litigation landscape shows ongoing risk and significant financial exposure for AI labs as publishers pursue enforcement.

  • Authors fear livelihood threats from AI tools trained on their works without consent, while some rulings offer cautious reassurance that training can be permissible under certain conditions.

  • Thomson Reuters v. Ross Intelligence illustrates that training content which directly competes with the original work is unlikely to be fair use.

  • Google’s case hinges on permission for internal AI training and whether publishers’ materials were used beyond intended limits.

Summary based on 7 sources


Get a daily email with more Tech stories

More Stories