← the record
WF-T84TAX

BREIN takes down AI dataset for copyright infringement

Stichting BREIN took down a large Dutch-language dataset used to train AI models after discovering it contained illegal copies of tens of thousands of books, news articles, and subtitles. The dataset was compiled from copyrighted material without permission. The maker of the dataset signed a declaration promising not to infringe and provided information about recipients. BREIN is investigating which AI models used the dataset.

Date it happened
2024-08-13
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from stichtingbrein.nl. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.