← the record
WF-V0LB9E

OpenAI deleted book datasets used to train GPT-3

Newly unsealed documents in the Authors Guild class-action lawsuit against OpenAI reveal that the startup deleted two datasets, 'books1' and 'books2,' which were used to train its GPT-3 model. The datasets allegedly contained over 100,000 published books and constituted 16% of GPT-3's training data. OpenAI says the datasets were deleted in mid-2022 due to nonuse, and the employees who created them are no longer with the company. The Authors Guild alleges copyright infringement, and the dispute over sealing the employees' names and dataset details continues.

Organisation involved
OpenAI
Product, system or model
GPT-3
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from businessinsider.com. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.