DiveFace dataset criticised for violating privacy, promoting harmful stereotyping
The DiveFace dataset was accused of committing legal violations and breac hing ethical norms about the practices of its developers. The dataset contains biometric data (facial images) of 24,000 individuals collected from Flickr without their consent, violating the privacy of those individuals whose personal data was repurposed for developing facial recognition technology, according to Exposing.ai. DiveFace was also discovered to be categorising individuals into broad and reductive ethnic groups (East Asian, Sub-Saharan and South Indian, and Caucasian), thereby oversimplifying human diversity and promoting harmful stereotyping. Furthermore, a significant portion of the images in DiveFace were found to be licensed under Creative Commons BY-NC-ND, which prohibits commercial use and derivations. The use of images in the DiveFace dataset by commercial entities violates the license terms, but DiveFace provides open access to its data without restrictions. System 🤖 DiveFace Operator: Developer: Aythami Morales, Julian Fierrez, Ruben Vera-Rodriguez, Ruben Tolosana Country: Global Sector: Research/academia; Technology Purpose: Train facial recognition systems Technology: Database/dataset; Facial recognition; Computer vision Issue: Bias/discrimination - race, ethnicity; Copyright; Privacy; Tr ansparency
- Date it happened
- 2021-01-01
- Product, system or model
- DiveFace
This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.