Big Tech companies used YouTube videos to train AI without consent
Proof News found that subtitles from 173,536 YouTube videos were used by companies including Anthropic, Nvidia, Apple, and Salesforce to train AI models. The dataset, called YouTube Subtitles, was created by EleutherAI and published in 2020. Creators were not aware and some have expressed frustration, calling it theft. The companies have acknowledged using the dataset but argue it was publicly available.
- Date it happened
- 2020-01-01
- Organisation involved
- Anthropic, Nvidia, Apple, Salesforce, Bloomberg, Databricks
- Product, system or model
- Claude, OpenELM
This incident was imported from proofnews.org. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.