OpenAI scrapes YouTube to train GPT-4
OpenAI quietly scraped and transcribed over one million hours of YouTube videos to train its GPT-4 large language model, raising questions about copyright and ethics. In an effort to source additional training data, Open AI created Whisper , a speech recognition system to transcribe over a million hours of YouTube videos. The resulting transcriptions were included in the training data for GPT-4 . Experts have demonstrated that the performance of large language models improves with increased amounts of training data. According to The New York Times , OpenAI knew its use of YouTube videos without permission was legally questionable but believed it to be fair use. But critics say it goes against Google’s terms which restrict automated access to YouTube videos and the use of its videos for independent applications outside of YouTube. Google is thought likely to have known about OpenAI’s activities, but appears not to have acted, possibly as it has also been accused of using YouTube videos to train its own AI models. YouTube creators are thought to face a variety of actual and potential harms as a result of OpenAI’s actions, including copyright violations, financial loss, and privacy abuse. System 🤖 ChatGPT GPT-4 Whisper Operator: OpenAI Developer: OpenAI Country: Global Sector: Media/entertainment/sports/arts Purpose: Generate text Technology: Chatbot; NLP/text analysis; Neural network; Deep learning; Machine learning; Reinforcement learning Issue: Copyright; Transparency
- Date it happened
- 2024-04-01
- Organisation involved
- OpenAI
- Product, system or model
- GPT-4; ChatGPT
This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.