Copyright watchdog takes down Dutch language AI training dataset
A large dataset of copyrighted books and news articles was removed from the internet after an enforcement action by Dutch copyright enforcement group BREIN. The dataset, which remains unnamed, contained information collected without permission from tens of thousands of Dutch language books, news sites and subtitles from numerous films and TV series, and we being offered for use in training AI models, notably large language models. It is unclear how widely this dataset may have already been used by AI companies. BREIN director Bastiaan van Ramshorst said they were trying to act preemptively to avoid future lawsuits. The dataset was seen to raise questions about the legality and ethics of using copyrighted material for AI training without permission . The European Union's AI Act requires AI firms to disclose the datasets used to train their models. Copyright law of the European Union The copyright law of the European Union is the copyright law applicable within the European Union . Copyright law is largely harmonized in the Union, although country to country differences exist. Source: Wikipedia 🔗 System 🤖 Unknown Operator: Developer: Country: Netherlands Sector: Media/entertainment/sports/arts Purpose: Train AI models Technology: Data base/dataset Issue: Copyright
- Date it happened
- 2024-08-01
This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.