Meta Allegedly Used Books3, a Dataset of 191,000 Pirated Books, to Train LLaMA AI
Meta and Bloomberg allegedly used Books3, a dataset containing 191,000 pirated books, to train their AI models, including LLaMA and BloombergGPT, without author consent. Lawsuits from authors such as Sarah Silverman and Michael Chabon claim this constitutes copyright infringement. Books3 includes works from major publishers like Penguin Random House and HarperCollins. Meta argues its AI outputs are not "substantially similar" to the original books, but legal challenges continue.
- Date it happened
- 2020-10-25
- Organisation involved
- Meta, EleutherAI, Bloomberg, Generative AI developers
- Product, system or model
- The Pile, LLaMA, hugging face, GPT-J, Books3, BloombergGPT, Bibliotik
This incident was imported from AI Incident Database and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.