← the record
WF-44QN82

OpenAI's Training Data for LLMs Allegedly Comprised of Copyrighted Books

Two authors alleged in a class action lawsuit OpenAI infringed authors' copyrights by incorporating illegal "shadow libraries" offering copyrighted books without permission in the training data of its generative LLMs, such as ChatGPT.

Date it happened
2018-06-11
Organisation involved
OpenAI
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from AI Incident Database and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.