Researchers extract training data from ChatGPT using repeat prompt
Researchers from Google DeepMind and other institutions developed an attack that extracts verbatim training data from OpenAI's ChatGPT. By prompting the model to repeat a word forever, they caused it to output memorized personal information, including email addresses and phone numbers. The attack demonstrates that alignment techniques can be circumvented, and the researchers estimate that a gigabyte of training data could be extracted.
- Organisation involved
- OpenAI
- Product, system or model
- ChatGPT
This incident was imported from not-just-memorization.github.io. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.