MIT apologizes and pulls offline dataset that taught AI to use slurs
MIT permanently removed its 80 Million Tiny Images dataset after researchers discovered it contained thousands of images labeled with racist and misogynistic slurs. The dataset, used to train AI systems for object detection, was built by automatically scraping images from the internet using a list of nouns from WordNet that included derogatory terms. MIT apologized and urged the community to delete copies of the dataset.
- Date it happened
- 2020-07-01
- Organisation involved
- Massachusetts Institute of Technology
- Product, system or model
- 80 Million Tiny Images
This incident was imported from theregister.com. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.