The record

Where automated decisions went wrong

Incidents gathered from public reporting around the world. Each one links to the articles it came from. None of it is a finding that anyone broke the law.

Reports people file about their own experience are not shown here and never will be without their agreement. Tell us what happened to you.

Clear

92 incidents closest to “Provenir Data Marketplace” · matched on meaning · public reporting

iRobot considers selling Roomba home mapping data to advertisers

iRobot's Roomba vacuum cleaners with mapping technology collect detailed floor plan data of users' homes. The company's CEO stated he is considering selling this data to companies like Amazon, Apple, or Google for advertising purposes. Robotics experts express concern about the lack of user consent and potential privacy implications.

Company involved
iRobot
AI system involved
Roomba 960/980

10 source articles · read the reporting →

WF-BQBMHB4 Aug 2023

WorldCoin suspended in Kenya over data security concerns

WorldCoin, a digital identification protocol using iris scans, was suspended by Kenyan regulators (ODPC and Communications Authority) over concerns about data security, consent, and oversight. The system had issued digital IDs and cryptocurrency tokens to over 350,000 Kenyans. Reports of hacked orb operators and iris scans traded on the dark web have also emerged.

Company involved
Tools for Humanity GmbH
AI system involved
WorldCoin

10 source articles · read the reporting →

LINAGORA closes Lucie 7B after user mockery

LINAGORA, a French open-source software company, launched a beta version of its large language model Lucie 7B. The model was intended to be a transparent and ethical alternative to big tech AI. However, after users tested it and highlighted its shortcomings, the model was mocked online. LINAGORA subsequently closed the platform to address the issues and collect more data.

Company involved
LINAGORA
AI system involved
Lucie 7B

6 source articles · read the reporting →

WF-UO2QO927 Jan 2025

OpenAI accuses DeepSeek of inappropriately using its data

OpenAI has accused Chinese AI company DeepSeek of inappropriately using data from its ChatGPT model to train DeepSeek's own large language model. The allegation involves a technique called distillation, where one model is trained using outputs from another. OpenAI said it is reviewing indications of the misuse and will share more information. DeepSeek has not responded to the accusation.

Company involved
DeepSeek
AI system involved
DeepSeek

6 source articles · read the reporting →

WF-L8981D29 Jan 2025

DeepSeek exposed user data via open ClickHouse database

Cloud security firm Wiz discovered a ClickHouse database belonging to DeepSeek that was open to the internet without authentication, containing over a million lines of logs with chat histories, secret keys and backend details. Wiz disclosed the breach to DeepSeek, which promptly locked down the database. The incident highlights security risks in rapidly deploying AI services.

Company involved
DeepSeek
AI system involved
DeepSeek-R1

5 source articles · read the reporting →

WF-G1A5LS1 Nov 2023

ETH Zurich study shows LLMs can infer Reddit users' personal data

Researchers at ETH Zurich conducted a study where nine large language models, including GPT-4, analysed Reddit users' posts and inferred personal attributes such as age, location, gender, and income with up to 85% accuracy. The study randomly selected 520 users and found that GPT-4 was most accurate, while LlaMA-2-7b was least. The researchers warn that people unknowingly reveal personal information online that LLMs can exploit.

Company involved
ETH Zurich
AI system involved
GPT-4, LlaMA-2-7b

4 source articles · read the reporting →

WF-2ABL3N30 Nov 2023

Bavarian police test Palantir data mining with real personal data

The Bavarian State Criminal Police Office (LKA) has been testing Palantir's data mining software, called VeRa, with real personal data for months. The Bavarian data protection commissioner only learned of the test through a media inquiry and has announced a review. The Interior Ministry claims the test is lawful under current law, but critics argue a legal basis is missing.

Company involved
Bayerisches Landeskriminalamt
AI system involved
VeRa

7 source articles · read the reporting →

N-Tech.lab's FindFace used to identify St Petersburg metro passengers without consent

Egor Tsvetkov photographed passengers on the St Petersburg metro without their permission and used N-Tech.lab's facial recognition service FindFace to match their faces to public Vkontakte profiles. He published the results in an art project called 'Your Face is Big Data', saying he wanted to show how 'digital narcissism' can lead to stalking. Privacy advocates said the project was ethically problematic because the subjects had not consented and their identities were exposed. FindFace had been launched by N-Tech.lab in February 2016.

Company involved
N-Tech.lab
AI system involved
FindFace

8 source articles · read the reporting →

Spanish police bust $20M AI-powered investment scam

Spanish law enforcement, collaborating with international authorities, dismantled a $20 million investment scam that used AI-driven algorithms to deceive individuals and organizations. Six suspects were detained and assets, including luxury cars and cryptocurrency, were seized. The article does not report any compensation for victims.

5 source articles · read the reporting →

Pravda network content found in Wikipedia and AI chatbot outputs

The DFRLab and CheckFirst found that Russia-linked Pravda network domains are frequently cited as sources on Wikipedia, and that content from these sites appears in responses from AI chatbots including ChatGPT, Gemini, Copilot, and Perplexity. The chatbots did not disclose the network's links to Russia. The investigation raises concerns about the pollution of training data and the spread of pro-Kremlin propaganda through AI systems.

AI system involved
ChatGPT, Gemini, Copilot, Perplexity

10 source articles · read the reporting →

WF-7YJTZ215 Feb 2024

University of Michigan halts vendor offering student data for AI training

The University of Michigan asked a vendor to stop work after a LinkedIn message offered to license student data for AI training for $25,000. The data came from past research studies and did not contain personal identifiers. The university stated that student data was never for sale and that the vendor had shared inaccurate information. The vendor was asked to halt their work.

Company involved
University of Michigan

5 source articles · read the reporting →

AAIP investigates Worldcoin's personal data processing in Argentina

The Argentine Agency for Access to Public Information (AAIP) has initiated an investigation into the data processing practices of Worldcoin in Argentina. The investigation focuses on the collection, storage, and use of biometric data, including facial and iris scans, carried out in several cities in exchange for financial compensation. The AAIP aims to verify compliance with the country's data protection law, Ley 25.326, regarding sensitive data handling.

Company involved
Worldcoin (Fundación Worldcoin)
AI system involved
Worldcoin

8 source articles · read the reporting →

WF-59AG1J1 Dec 2021

Worldcoin collected biometric data from poor villagers in Indonesia without informed consent

Worldcoin, a cryptocurrency startup, recruited users in developing countries by offering free cash in exchange for iris scans. The company used deceptive marketing, collected more personal data than acknowledged, and failed to obtain meaningful informed consent. Many users received worthless tokens instead of promised money. The company acknowledged some friction but continued its operations.

Company involved
Worldcoin
AI system involved
chrome orb

5 source articles · read the reporting →

WF-6BTWTA8 Mar 2024

Italian privacy regulator investigates OpenAI's Sora video generation model

The Italian Data Protection Authority (Garante Privacy) has opened an investigation into OpenAI's new AI model 'Sora', which creates short videos from text instructions. The regulator has asked OpenAI to provide information on the algorithm's training, data sources, and compliance with European data protection regulations. OpenAI must respond within 20 days.

Company involved
OpenAI
AI system involved
Sora

7 source articles · read the reporting →

CJEU rules Dun & Bradstreet must explain automated credit decisions under GDPR

A customer was refused a mobile phone contract because of an automated credit assessment by Dun & Bradstreet Austria. The customer took the case to court, which found that Dun & Bradstreet had infringed the GDPR by failing to provide meaningful information about the logic involved. The CJEU ruled that data controllers must explain automated decisions and that trade secrets cannot automatically override the right of access.

Company involved
Dun & Bradstreet Austria GmbH

7 source articles · read the reporting →

WF-NL6UK61 Jan 2020

Big Tech companies used YouTube videos to train AI without consent

Proof News found that subtitles from 173,536 YouTube videos were used by companies including Anthropic, Nvidia, Apple, and Salesforce to train AI models. The dataset, called YouTube Subtitles, was created by EleutherAI and published in 2020. Creators were not aware and some have expressed frustration, calling it theft. The companies have acknowledged using the dataset but argue it was publicly available.

Company involved
Anthropic, Nvidia, Apple, Salesforce, Bloomberg, Databricks
AI system involved
Claude, OpenELM

10 source articles · read the reporting →

Delta uses AI from Fetcherr for domestic ticket pricing

Delta Air Lines is using generative AI from Fetcherr to determine some domestic flight prices, currently covering 3% of its network with plans to reach 20% by end of 2025. Democratic senators expressed concern that the AI could be used for individualized pricing based on personal data, leading to higher fares. Delta denies using personal data in pricing and states it complies with regulations. No actual harm has been reported.

Company involved
Delta Air Lines
AI system involved
Fetcherr

8 source articles · read the reporting →

News/Media Alliance study finds unauthorised use of publisher content to train AI

The News/Media Alliance alleges that generative AI developers have copied and used publishers' content without authorisation to train large language models. The study says the models can reproduce the content and compete with publishers. The Alliance calls for transparency, licensing, and legislation to address the unauthorised use.

10 source articles · read the reporting →

WF-T84TAX13 Aug 2024

BREIN takes down AI dataset for copyright infringement

Stichting BREIN took down a large Dutch-language dataset used to train AI models after discovering it contained illegal copies of tens of thousands of books, news articles, and subtitles. The dataset was compiled from copyrighted material without permission. The maker of the dataset signed a declaration promising not to infringe and provided information about recipients. BREIN is investigating which AI models used the dataset.

8 source articles · read the reporting →

WF-PECTZC3 Sep 2024

Dutch DPA fines Clearview AI €30.5 million for GDPR violations

The Dutch Data Protection Authority fined Clearview AI €30.5 million for creating an illegal database of biometric codes scraped from social media without consent. The regulator accused the company of failing to inform individuals about how their data is used and continuing violations after investigation. Clearview AI denies the allegations, stating it has no presence in the Netherlands or EU and that the fine is unenforceable.

Company involved
Clearview AI
AI system involved
Clearview AI facial recognition database

7 source articles · read the reporting →

Publishers sue AI startup Cohere over alleged copyright infringement

A consortium of 14 publishers including Condé Nast, The Atlantic, and Forbes filed a lawsuit against Cohere, alleging that the generative AI startup engaged in massive, systematic copyright infringement by using at least 4,000 copyrighted works to train its AI models and display large portions of articles, harming referral traffic. Cohere denied the allegations, calling the lawsuit misguided and frivolous.

Company involved
Cohere
AI system involved
Cohere's AI models

8 source articles · read the reporting →

WF-EBIPSN1 Dec 2025

Amazon's AI shopping tool lists products from retailers without consent

Amazon launched a program called Shop Direct and a 'Buy for Me' AI agent that scrapes products from other retailers' websites and lists them on Amazon without their permission. Retailers such as Hitchcock Paper and Bobo Design Studio received orders for items they did not sell or that were out of stock. Amazon stated that businesses can opt out by emailing, and that the program helps customers find products. Some retailers have accused Amazon of exploiting their businesses.

Company involved
Amazon
AI system involved
Shop Direct / Buy for Me

5 source articles · read the reporting →

Italian DPA fines Municipality of Trento over AI surveillance projects

The Italian data protection authority (Garante) fined the Municipality of Trento €50,000 for two research projects, Marvel and Protector, that used AI to analyze video, audio, and social media data for public security purposes. The projects involved automated detection of risk events from surveillance cameras and microphones in public spaces, as well as monitoring social media for hate speech. The Garante found multiple violations of privacy law, including lack of a valid legal basis, insufficient anonymization, failure to conduct a data protection impact assessment, and inadequate transparency. The municipality is required to delete the unlawfully processed data.

Company involved
Comune di Trento
AI system involved
Marvel and Protector

9 source articles · read the reporting →

WF-PCWX3U1 Jan 2024

OpenAI's GPT Store hosts copyright-infringing chatbots

Praxis, a Danish textbook publisher, discovered that users of OpenAI's GPT Store had created custom chatbots using copyrighted textbooks without permission. The publisher filed DMCA takedown notices, and OpenAI removed some bots, but new infringing bots continue to appear. Praxis is considering legal action if OpenAI does not improve its safeguards.

Company involved
OpenAI
AI system involved
GPT Store

3 source articles · read the reporting →

← Newerpage 3 of 4Older →