← the record
WF-7VGPYK

OpenAI transcribed YouTube videos to train GPT-4 without permission

OpenAI used its Whisper transcription model to transcribe over a million hours of YouTube videos, according to a New York Times report. The company allegedly used the transcripts to train GPT-4 despite knowing the practice was legally questionable. Google, which owns YouTube, said it prohibits unauthorized scraping of its content. OpenAI has said it believes its use of the data constitutes fair use.

Date it happened
2021-01-01
Organisation involved
OpenAI
Product, system or model
Whisper, GPT-4
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from theverge.com. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.