DiveFace face-recognition dataset reuses Flickr photos despite licence restrictions
DiveFace is a face-recognition dataset published in 2019 with 139,677 images of about 24,000 people. The authors say the images came from Flickr via the MegaFace and YFCC100M datasets and were automatically labelled by gender and ethnicity. Exposing.ai found that many of the photos carry Creative Commons licences which prohibit derivative use, suggesting the dataset may have been assembled without proper permission. The dataset remains available from the authors' GitHub page.
- Date it happened
- 2019-01-01
- Product, system or model
- DiveFace
This incident was imported from exposing.ai. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.