← the record
AIAAIC-1531

Study: Hate content increases 12 percent as LAION dataset size increases

A comparative audit of two datasets, LAION-400M and LAION-2B, revealed that as the dataset scale increases, hate content also increases by nearly 12 percent . In a recent study titled “Into the LAION’s Den: Investigating Hate in Multimodal Datasets,” researchers examined the impact of scaling datasets on hateful content by comparing two datasets: LAION-400M and LAION-2B. The results of the audit revealed that hate content increased by nearly 12 percent as the dataset size grew - an increase measured qualitatively and quantitatively using the Hate Content Rate (HCR) metric . The finding highlighted the consequences of data scaling in vision-language datasets . System 🤖 LAION-400M Operator: Alphabet/Google; Prisma Labs; Stability AI Developer: LA ION Country: Germany Sector: Technology Purpose: Train large language models Technology: Database/dataset; Neural network; Deep learning; Machine learning Issue: Safety

Date it happened
2023-01-01
Organisation involved
Google; Prisma Labs; Stability AI
Product, system or model
LAION-400M
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.