Study: Hate content increases 12 percent as LAION dataset size increases
A comparative audit of two datasets, LAION-400M and LAION-2B, revealed that as the dataset scale increases, hate content also increases by nearly 12 percent . In a recent study titled “Into the LAION’s Den: Investigating Hate in Multimodal Datasets,” researchers examined the impact of scaling datasets on hateful content by comparing two datasets: LAION-400M and LAION-2B. The results of the audit revealed that hate content increased by nearly 12 percent as the dataset size grew - an increase measured qualitatively and quantitatively using the Hate Content Rate (HCR) metric . The finding highlighted the consequences of data scaling in vision-language datasets . System 🤖 LAION-400M Operator: Alphabet/Google; Prisma Labs; Stability AI Developer: LA ION Country: Germany Sector: Technology Purpose: Train large language models Technology: Database/dataset; Neural network; Deep learning; Machine learning Issue: Safety
- Date it happened
- 2023-01-01
- Organisation involved
- Google; Prisma Labs; Stability AI
- Product, system or model
- LAION-400M
This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.