← the record
WF-JLTHZQ

Meta Allegedly Used Books3, a Dataset of 191,000 Pirated Books, to Train LLaMA AI

Meta and Bloomberg allegedly used Books3, a dataset containing 191,000 pirated books, to train their AI models, including LLaMA and BloombergGPT, without author consent. Lawsuits from authors such as Sarah Silverman and Michael Chabon claim this constitutes copyright infringement. Books3 includes works from major publishers like Penguin Random House and HarperCollins. Meta argues its AI outputs are not "substantially similar" to the original books, but legal challenges continue.

Date it happened
2020-10-25
Organisation involved
Meta, EleutherAI, Bloomberg, Generative AI developers
Product, system or model
The Pile, LLaMA, hugging face, GPT-J, Books3, BloombergGPT, Bibliotik
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from AI Incident Database and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.