The Pile dataset
The Pile is a 886 GiB, open-source dataset of English-language text created to help train large language models (LLMs). Developed by EleutherAI and publicly released in December 2020, The Pile consists of 22 smaller datasets, including Books3 , BookCorpus and YouTube Subtitles , and 14 new ones . Initially developed to train EleutherAI's GPT-Neo models, The Pile has been used to train many other models, including Microsoft's Megatron-Turing Natural Language Generation, Meta AI's Open Pre-trained Transformers, LL a MA , and Galactica , Stanford University's BioMedLM 2.7B, the B e ijing Academy of Artificial Intelligence 's Chinese-Transformer-XL, Yandex 's YaLM 100B, and Apple's OpenELM. Dataset 🤖 The Pile dataset 🔗 Released: 2020 Developer: EleutherAI Purpose: Train large language models Type: Database/dat aset Technique: Generative AI; Large language model; Machine learning Transparency, accountability 🙈 The Pile is seen to suffer from multiple transparency and accountability limitations: Inadequate documentation . The dataset has not been thoroughly documented by its creators, making it difficult for researchers to fully understand and address potential issues. Lack of clear filtering mechanisms. The Pile does not appear to have implemented extensive processes for filtering content deemed toxic or private, leaving much of this responsibility to individual researchers using the dataset. Privacy consent. Several datasets included in The Pile were collected without the explicit consent of the individuals whose data is included. This is particularly concerning for datasets like the Enron emails, where individuals had no opportunity to consent to their inclusion. Copyright compliance. Some components of The Pile may contain copyrighted material that was not collected or distributed in compliance with terms of service agreements. This raises legal and ethical questions about the use of such data. Limited accountability. There appears to be a lack of established mechanisms for reporting and remedying faults or biases discovered in the dataset. Resources 📃 Datasheet for The Pile The Pile: An 800GB Dataset of Diverse Text for Language Modeling Derivatives, applications 🈸 Galactica LLaMA Risks, harms 🛑 The Pile dataset has been accused of copyright and privacy abuse, and of enabling the creation and deployment of biased, unsafe and unethical AI models. Incidents, issues 🔥 Anthropic "destructively" scans millions of books to train AI models Books3 dataset shut down after legal notice from Danish anti-piracy group Nvidia sued for training NeMo on authors' copyrighted works Mike Huckabee books used to train language models without consent Sarah Silverman sues OpenAI for violating copyright OpenAI deleted training datasets believed to contain copyrighted books 17 authors sue OpenAI for 'systematic mass-scale copyright infringement'
- Organisation involved
- Apple; Beijing Academy of Artificial Intelligence; EleutherAI; Meta; Microsoft; Stanford University; Yandex
- Product, system or model
- The Pile
This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.