IARPA Janus Benchmark C dataset used faces from YouTube without consent
The IARPA Janus Benchmark C (IJB-C) dataset, published in 2017, contains 21,294 images and names of 3,531 individuals scraped from YouTube, Flickr, and Wikimedia without their consent. The dataset was created for face recognition benchmarking to improve intelligence analysis. Jillian York, a digital rights activist, had 41 frames of her face taken from a YouTube video without her knowledge or permission. The dataset violates YouTube's terms of service and raises privacy concerns.
- Date it happened
- 2017-01-01
- Organisation involved
- IARPA (Intelligence Advanced Research Projects Activity)
- Product, system or model
- IARPA Janus Benchmark C (IJB-C)
This incident was imported from exposing.ai. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.