LFW dataset discards the privacy rights of internet users
Prominent dataset Labeled Faces in the Wild (LFW) quietly scraped Google, Flickr, YouTube and other online photo libraries, discarding the privacy rights of photo owners and subjects. In a paper examining over 130 facial-recognition data sets compiled over 43 years, researchers Deborah Raji and Genevieve Fried singled out the LFW dataset as being the first for which 'wild' images were scraped from the internet . According to the Technology Review , the dataset 'opened the floodgates to data collection through web search , with r esearchers starting to download images directly from Google, Flickr, and Yahoo without concern for consent.' Earlier, LFW had been found to be highly skewed towards a very small subset of people, specifically white male faces, and contain ed 'a significant number of duplicate or nearly-duplicate images and mislabeled images.' The finding persuaded LFW's creators to acknow ledge the dataset's limitations. System 🤖 Labeled Faces in the Wild Operator: Developer: University of Massachussets, Amherst Country: USA Sector: Research/academia; Technology Purpose: Train facial recognition systems Technology: Dataset; Computer vision; Deep learning; Facial recognition; Facial detection; Facial analysis; Machine learning; Neural network; Pattern recognition Issue: Bias/discrimination - race, ethnicity, gender; Ethics/values; Privacy; Tr ansparency
- Date it happened
- 2021-01-01
- Product, system or model
- Labeled Faces in the Wild (LFW)
This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.