GPT-2 Reportedly Reproduced Personal Data from Its Training Data
OpenAI's GPT-2 reportedly memorized and reproduced portions of its training data, including personal information such as names, email addresses, social media handles, and phone numbers. Researchers raised concerns that large language models could expose private or sensitive information when trained on web-scale datasets containing personal data.
- Date it happened
- 2019-02-14
- Organisation involved
- OpenAI
- Product, system or model
- GPT-2, Large language models, Chatbots
This incident was imported from AI Incident Database and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.