← the record
AIAAIC-1590

YouTube Subtitles dataset

YouTube Subtitles is a dataset that comprises subtitles from 173,536 YouTube videos taken from more than 48,000 channels, including Khan Academy, MIT, Harvard University, the Wall Street Journal , NPR , the BBC , PewDiePie and Mr Beast. The subtitles are often presented alongside translations into languages such as Japanese, German and Arabic. Released in 2020 as part of The Pile dataset, YouTube subtitles has been used by multiple technology companies, including Anthropic, Nvidia, Apple, Bloomberg and Salesforce. Dataset 🤖 Dataset 🔗 Released: 2020 Developer: EleutherAI Purpose: Train AI models Type: Database/dat aset Technique: Machine learning Transparency, accountability 🙈 Data sources. AI companies have not been open about their data used to train their AI models, including YouTube Subtitles. Accountability. EleutherAI declined to discuss allegations that YouTube videos used without permission to create YouTube Subtitles formed part its Pile dataset. Risks, harms 🛑 The YouTube Subtitles dataset has been criticised for transcribing the output of video creators without their explicit permission, thereby potentially violating their copyright. Incidents, issues 🔥 July 2024. Apple, Nvidia, Anthropic used thousands of YouTube videos without permission to train AI models

Organisation involved
Anthropic; Apple; Nvidia; Salesforce
Product, system or model
YouTube Subtitles
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.