← the record
AIAAIC-1427

Nvidia sued for training NeMo on authors' copyrighted works

GPU chip provider NVIDIa has been sued by 3 authors accusing it of training it's NeMo AI models on copyrighted books. A uthors Brian Keene, Abdi Nazemian and Stewart O'Nan submitted a class action lawsuit against N VIDIA for copyright infringement, saying their works were part of th e Books3 dataset and were trained on NeMO generative AI platform without their permission. The Books3 dataset, the lawsuit argued, copied "all of Bibliotek" - a so-called shadow library of approximately 196,640 pirated books that had earlier been available as part of The Pile - a larger dataset - through AI community Hugging Face. The Pile was later removed from Hugging Face in the wake of a copyright complaint. The authors want compensation for their creative labour and the destruction of all copies of the Books3 dataset, and argue that NVIDIA’s October 2023 takedown of the NeMo AI platform was an implicit admission of its guilt. The case highlighted ongoing copyright clashes between the AI industry and creative communities, with transparency and infringement claims at the forefront . ➖ October 2023. Nvidia withdrew the NeMo platform and acknowledged the model had been trained on a dataset containing "approximately " 196,640 books. The Books3 dataset contains the same number of books. Fair use Fair use is a doctrine in United States law that permits limited use of copyrighted material without having to first acquire permission from the copyright holder. Source: Wikipedia 🔗 System 🤖 Nvidia NeMo 🔗 Books3 Operator: Nvidia Developer: Nvidia Country: USA Sector: Media/entertainment/sports/arts Purpose: Train and deploy custom LLMs Technology: Generative AI; Machine learning; Neural network; Deep learning; NLP/text analysis Issue: Accountability; Copyright; Ethics/values; Transparency Regulation 👩🏼‍⚖️ Digital Millennium Copyright Act (DMCA)

Date it happened
2024-01-01
Organisation involved
Nvidia
Product, system or model
NeMo
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.