← the record
WF-AW8J1O

Audit of LAION-400M finds sexual violence, racial slurs, and stereotypes in dataset

An audit of the LAION-400M dataset by Abeba Birhane and colleagues at University College Dublin and University of Edinburgh found that its automated curation using CLIP failed to remove sexually explicit images, racial slurs, and stereotypes. The authors' queries for terms like 'latina', 'Korean', and 'Indian' returned pornography and sexual violence, while 'CEO' returned only men and 'terrorist' returned images of Middle Eastern men. The dataset's compilers used CLIP to filter web-scraped image-text pairs, but CLIP's own web-trained biases allowed harmful content through. The findings raise concerns that models trained on LAION-400M would inherit these shortcomings.

Date it happened
2021-09-01
Organisation involved
LAION-400M team
Product, system or model
LAION-400M
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from info.deeplearning.ai. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.

Audit of LAION-400M finds sexual violence, racial slurs, and stereotypes in dataset — Wayward Fowl