The record

Where automated decisions went wrong

Incidents gathered from public reporting around the world. Each one links to the articles it came from. None of it is a finding that anyone broke the law.

Reports people file about their own experience are not shown here and never will be without their agreement. Tell us what happened to you.

Clear

116 incidents closest to “Provenir Data Marketplace” · matched on meaning · public reporting

Pravda network content found in Wikipedia and AI chatbot outputs

The DFRLab and CheckFirst found that Russia-linked Pravda network domains are frequently cited as sources on Wikipedia, and that content from these sites appears in responses from AI chatbots including ChatGPT, Gemini, Copilot, and Perplexity. The chatbots did not disclose the network's links to Russia. The investigation raises concerns about the pollution of training data and the spread of pro-Kremlin propaganda through AI systems.

AI system involved
ChatGPT, Gemini, Copilot, Perplexity

10 source articles · read the reporting →

WF-7YJTZ215 Feb 2024

University of Michigan halts vendor offering student data for AI training

The University of Michigan asked a vendor to stop work after a LinkedIn message offered to license student data for AI training for $25,000. The data came from past research studies and did not contain personal identifiers. The university stated that student data was never for sale and that the vendor had shared inaccurate information. The vendor was asked to halt their work.

Company involved
University of Michigan

5 source articles · read the reporting →

AAIP investigates Worldcoin's personal data processing in Argentina

The Argentine Agency for Access to Public Information (AAIP) has initiated an investigation into the data processing practices of Worldcoin in Argentina. The investigation focuses on the collection, storage, and use of biometric data, including facial and iris scans, carried out in several cities in exchange for financial compensation. The AAIP aims to verify compliance with the country's data protection law, Ley 25.326, regarding sensitive data handling.

Company involved
Worldcoin (Fundación Worldcoin)
AI system involved
Worldcoin

8 source articles · read the reporting →

WF-59AG1J1 Dec 2021

Worldcoin collected biometric data from poor villagers in Indonesia without informed consent

Worldcoin, a cryptocurrency startup, recruited users in developing countries by offering free cash in exchange for iris scans. The company used deceptive marketing, collected more personal data than acknowledged, and failed to obtain meaningful informed consent. Many users received worthless tokens instead of promised money. The company acknowledged some friction but continued its operations.

Company involved
Worldcoin
AI system involved
chrome orb

5 source articles · read the reporting →

WF-6BTWTA8 Mar 2024

Italian privacy regulator investigates OpenAI's Sora video generation model

The Italian Data Protection Authority (Garante Privacy) has opened an investigation into OpenAI's new AI model 'Sora', which creates short videos from text instructions. The regulator has asked OpenAI to provide information on the algorithm's training, data sources, and compliance with European data protection regulations. OpenAI must respond within 20 days.

Company involved
OpenAI
AI system involved
Sora

7 source articles · read the reporting →

CJEU rules Dun & Bradstreet must explain automated credit decisions under GDPR

A customer was refused a mobile phone contract because of an automated credit assessment by Dun & Bradstreet Austria. The customer took the case to court, which found that Dun & Bradstreet had infringed the GDPR by failing to provide meaningful information about the logic involved. The CJEU ruled that data controllers must explain automated decisions and that trade secrets cannot automatically override the right of access.

Company involved
Dun & Bradstreet Austria GmbH

7 source articles · read the reporting →

WF-GEPFGZ1 Dec 2023

OpenDream AI art site allowed users to generate child sexual abuse material

OpenDream, an AI image generation platform, allowed users to generate and publicly display child sexual abuse material (CSAM) and non-consensual deepfakes from at least December 2023 until July 2024. The platform, operated by CBM Media Pte Ltd in Singapore, offered paid plans with NSFW prompts and models. Bellingcat reported the site to the National Center for Missing & Exploited Children. After Bellingcat's inquiry, the CSAM was removed from the site and search engines, and Google terminated OpenDream's AdSense account.

Company involved
CBM Media Pte Ltd
AI system involved
OpenDream

3 source articles · read the reporting →

WF-NL6UK61 Jan 2020

Big Tech companies used YouTube videos to train AI without consent

Proof News found that subtitles from 173,536 YouTube videos were used by companies including Anthropic, Nvidia, Apple, and Salesforce to train AI models. The dataset, called YouTube Subtitles, was created by EleutherAI and published in 2020. Creators were not aware and some have expressed frustration, calling it theft. The companies have acknowledged using the dataset but argue it was publicly available.

Company involved
Anthropic, Nvidia, Apple, Salesforce, Bloomberg, Databricks
AI system involved
Claude, OpenELM

10 source articles · read the reporting →

WF-C7X63S22 Jul 2024

Condé Nast accuses Perplexity of plagiarism in cease-and-desist letter

Condé Nast, the media conglomerate, sent a cease-and-desist letter to AI search startup Perplexity, accusing it of plagiarism for using content from its publications in AI-generated responses without permission. The letter demands that Perplexity stop using the content. Perplexity has been criticized for ignoring robots.txt and scraping content. The incident highlights ongoing tensions between publishers and AI companies over unauthorized use of content.

Company involved
Perplexity
AI system involved
Perplexity

4 source articles · read the reporting →

Delta uses AI from Fetcherr for domestic ticket pricing

Delta Air Lines is using generative AI from Fetcherr to determine some domestic flight prices, currently covering 3% of its network with plans to reach 20% by end of 2025. Democratic senators expressed concern that the AI could be used for individualized pricing based on personal data, leading to higher fares. Delta denies using personal data in pricing and states it complies with regulations. No actual harm has been reported.

Company involved
Delta Air Lines
AI system involved
Fetcherr

8 source articles · read the reporting →

News/Media Alliance study finds unauthorised use of publisher content to train AI

The News/Media Alliance alleges that generative AI developers have copied and used publishers' content without authorisation to train large language models. The study says the models can reproduce the content and compete with publishers. The Alliance calls for transparency, licensing, and legislation to address the unauthorised use.

10 source articles · read the reporting →

WF-T84TAX13 Aug 2024

BREIN takes down AI dataset for copyright infringement

Stichting BREIN took down a large Dutch-language dataset used to train AI models after discovering it contained illegal copies of tens of thousands of books, news articles, and subtitles. The dataset was compiled from copyrighted material without permission. The maker of the dataset signed a declaration promising not to infringe and provided information about recipients. BREIN is investigating which AI models used the dataset.

8 source articles · read the reporting →

WF-PECTZC3 Sep 2024

Dutch DPA fines Clearview AI €30.5 million for GDPR violations

The Dutch Data Protection Authority fined Clearview AI €30.5 million for creating an illegal database of biometric codes scraped from social media without consent. The regulator accused the company of failing to inform individuals about how their data is used and continuing violations after investigation. Clearview AI denies the allegations, stating it has no presence in the Netherlands or EU and that the fine is unenforceable.

Company involved
Clearview AI
AI system involved
Clearview AI facial recognition database

7 source articles · read the reporting →

Publishers sue AI startup Cohere over alleged copyright infringement

A consortium of 14 publishers including Condé Nast, The Atlantic, and Forbes filed a lawsuit against Cohere, alleging that the generative AI startup engaged in massive, systematic copyright infringement by using at least 4,000 copyrighted works to train its AI models and display large portions of articles, harming referral traffic. Cohere denied the allegations, calling the lawsuit misguided and frivolous.

Company involved
Cohere
AI system involved
Cohere's AI models

8 source articles · read the reporting →

WF-EBIPSN1 Dec 2025

Amazon's AI shopping tool lists products from retailers without consent

Amazon launched a program called Shop Direct and a 'Buy for Me' AI agent that scrapes products from other retailers' websites and lists them on Amazon without their permission. Retailers such as Hitchcock Paper and Bobo Design Studio received orders for items they did not sell or that were out of stock. Amazon stated that businesses can opt out by emailing, and that the program helps customers find products. Some retailers have accused Amazon of exploiting their businesses.

Company involved
Amazon
AI system involved
Shop Direct / Buy for Me

5 source articles · read the reporting →

Italian DPA fines Municipality of Trento over AI surveillance projects

The Italian data protection authority (Garante) fined the Municipality of Trento €50,000 for two research projects, Marvel and Protector, that used AI to analyze video, audio, and social media data for public security purposes. The projects involved automated detection of risk events from surveillance cameras and microphones in public spaces, as well as monitoring social media for hate speech. The Garante found multiple violations of privacy law, including lack of a valid legal basis, insufficient anonymization, failure to conduct a data protection impact assessment, and inadequate transparency. The municipality is required to delete the unlawfully processed data.

Company involved
Comune di Trento
AI system involved
Marvel and Protector

9 source articles · read the reporting →

WF-4Q3RDL2 Feb 2023

Italian Data Protection Authority Blocks Replika Chatbot Over Risks to Minors

On February 2, 2023, the Italian Data Protection Authority (Garante) issued an urgent order blocking the AI chatbot Replika from processing personal data of Italian users. The Garante found that Replika lacked effective age verification, allowing minors to potentially receive inappropriate content including sex-related replies, and that its privacy policy violated GDPR transparency requirements. The U.S.-based controller was given 20 days to report on compliance measures and may challenge the order within 60 days.

AI system involved
Replika

1 source article · read the reporting →

WF-R41UQV23 Feb 2012

Palantir secretly tested predictive policing in New Orleans

Beginning in 2012, Palantir Technologies secretly partnered with the New Orleans Police Department to deploy a predictive policing system. The programme analysed gang affiliations, social media and criminal histories to forecast individuals’ likelihood of committing or becoming victims of violence, operating without public knowledge or city council oversight. Researchers and law enforcement officials raised concerns about systemic bias and civil liberties. As of 2018, the city and Palantir had not disclosed the programme’s status.

Company involved
New Orleans Police Department

1 source article · read the reporting →

WF-PCWX3U1 Jan 2024

OpenAI's GPT Store hosts copyright-infringing chatbots

Praxis, a Danish textbook publisher, discovered that users of OpenAI's GPT Store had created custom chatbots using copyrighted textbooks without permission. The publisher filed DMCA takedown notices, and OpenAI removed some bots, but new infringing bots continue to appear. Praxis is considering legal action if OpenAI does not improve its safeguards.

Company involved
OpenAI
AI system involved
GPT Store

3 source articles · read the reporting →

Adobe Firefly trained on thousands of Midjourney images, Bloomberg reports

Bloomberg has reported that Adobe's Firefly image generator was trained using thousands of images from competitor Midjourney. Adobe says these made up about 5% of the training data and were part of the Adobe Stock library. The company has marketed Firefly as ethically trained and offered enterprise customers indemnity against copyright claims. Adobe responded that all Adobe Stock images undergo moderation, but the report has raised questions about Firefly's copyright safety.

Company involved
Adobe
AI system involved
Firefly

7 source articles · read the reporting →

Clearview AI facial recognition app scrapes billions of images and is used by police

Clearview AI, a secretive start-up, built a facial recognition app that matches photos to a database of more than three billion images scraped from social media and websites. More than 600 law enforcement agencies, including the FBI and Department of Homeland Security, have used the tool to identify suspects in crimes such as shoplifting, identity theft and murder. The company monitored officers who ran a reporter's photo through the app, and its founder acknowledged designing an augmented-reality prototype but said there were no plans to release it. Critics warned the tool could end anonymity and enable misuse.

Company involved
Clearview AI
AI system involved
Clearview AI facial recognition app

2 source articles · read the reporting →

WF-AYLKD31 Dec 2020

NHS faces legal action over Palantir data contract extension

The NHS is being taken to court by campaign group Open Democracy over its contract with data firm Palantir. The legal action alleges that the extension of Palantir's involvement in analysing NHS patient data for pandemic response and beyond lacked a proper Data Protection Impact Assessment. The contract, initially an emergency response, was extended for two years at a cost of £23.5m. The case is pending.

Company involved
NHS
AI system involved
Palantir data analysis platform

10 source articles · read the reporting →

WF-NWB6EE1 Dec 2024

OpenAI's Sora video generator trained on copyrighted content without consent, tests suggest

Tests by The Washington Post suggest that OpenAI's video generation tool Sora was trained on videos from Netflix, TikTok, YouTube and other sources without permission. OpenAI has not disclosed its training data for Sora, saying only that it used publicly available and licensed data. The tool can closely replicate copyrighted characters, logos and scenes, raising concerns about copyright infringement. OpenAI has not faced a lawsuit specifically over Sora's training data but is fighting other copyright suits.

Company involved
OpenAI
AI system involved
Sora

3 source articles · read the reporting →

WF-NTJTJM27 May 2021

Privacy International challenges Clearview AI's facial recognition database in Europe

Privacy International filed complaints against Clearview AI with five European data protection authorities in May 2021, alleging that the company's scraping of facial images from the web and building a biometric database without consent violates data protection laws. The regulators in the UK, France, Italy, Greece, and Austria have since found Clearview's practices unlawful, imposed fines, and ordered deletion of data. Clearview has appealed the UK fine, and the case is ongoing.

Company involved
Clearview AI
AI system involved
Clearview

10 source articles · read the reporting →

← Newerpage 4 of 5Older →