The record

Where automated decisions went wrong

Incidents gathered from public reporting around the world. Each one links to the articles it came from. None of it is a finding that anyone broke the law.

Reports people file about their own experience are not shown here and never will be without their agreement. Tell us what happened to you.

Clear

92 incidents closest to “Duolingo English Test” · matched on meaning · public reporting

Tourists rescued after following ChatGPT route in Tatra mountains

Three Polish tourists used ChatGPT to plan a hiking route from Hala Gąsienicowa to Dolina Pięciu Stawów in the Tatra mountains. The AI recommended a difficult route through Przełęcz Krzyżne, which proved too dangerous in winter conditions with limited visibility and icy rocks. The tourists called the TOPR rescue service and were safely brought down. A TOPR rescuer advised against relying on ChatGPT or Google Maps for mountain route planning.

Company involved
OpenAI
AI system involved
ChatGPT

2 source articles · read the reporting →

WF-QLNRBN31 Jan 2023

ElevenLabs voice cloning tool used to create deepfake celebrity audio clips

Speech AI startup ElevenLabs launched a beta voice cloning tool. Within days, users on 4chan posted deepfake audio clips featuring voices resembling celebrities like Emma Watson reading offensive material. ElevenLabs acknowledged the misuse and said it is considering additional safeguards.

Company involved
ElevenLabs
AI system involved
ElevenLabs platform

10 source articles · read the reporting →

WF-3UCXWE1 Jul 2023

Study finds Midjourney, DALL-E 2, Stable Diffusion accept over 85% of fake news prompts

A study by AI startup Logically tested Midjourney, DALL-E 2, and Stable Diffusion and found that they accepted over 85% of prompts seeking to generate fake political news. The systems generated images of ballot stuffing, small boat arrivals, and explosions. Logically warned that the lack of moderation could pose threats to upcoming elections. Stability AI responded by stating its ethical use license and measures to prevent misuse.

Company involved
Midjourney, OpenAI, Stability AI
AI system involved
Midjourney, DALL-E 2, Stable Diffusion

8 source articles · read the reporting →

Answer.AI tests Devin and reports 14 failures in 20 tasks

Answer.AI's team tested Devin, an autonomous AI coding assistant, on 20 real-world tasks over a month. Devin succeeded in only 3 tasks, failed 14, and was inconclusive in 3. The team found Devin often produced overly complex or hallucinated solutions and could not recognize fundamental blockers. They ultimately decided to stick with tools that allow more human control.

AI system involved
Devin

5 source articles · read the reporting →

LINAGORA closes Lucie 7B after user mockery

LINAGORA, a French open-source software company, launched a beta version of its large language model Lucie 7B. The model was intended to be a transparent and ethical alternative to big tech AI. However, after users tested it and highlighted its shortcomings, the model was mocked online. LINAGORA subsequently closed the platform to address the issues and collect more data.

Company involved
LINAGORA
AI system involved
Lucie 7B

6 source articles · read the reporting →

WF-MZCD6720 Sep 2023

Polish DPO investigates OpenAI over ChatGPT false data and lack of transparency

The Polish data protection authority (UODO) is investigating a complaint against OpenAI concerning ChatGPT. The complainant alleges that ChatGPT generated false information about him, and that OpenAI failed to correct it or disclose what data it holds, violating GDPR principles of lawfulness, fairness and transparency. The complainant also claims OpenAI did not fulfil its information obligations under Article 12 and Article 5(1)(a) GDPR. UODO has stated it will examine the systemic compliance of OpenAI's data processing with European data protection law.

Company involved
OpenAI
AI system involved
ChatGPT

9 source articles · read the reporting →

Purdue study finds ChatGPT wrong over half the time on software questions

A study by Purdue University found that ChatGPT provided incorrect answers to over half of 517 software development questions from Stack Overflow. Despite the errors, 34% of users preferred ChatGPT's answers over human responses. The study warns that relying on ChatGPT for coding could jeopardize programmers' professional reputations.

Company involved
OpenAI
AI system involved
ChatGPT

9 source articles · read the reporting →

WF-G1A5LS1 Nov 2023

ETH Zurich study shows LLMs can infer Reddit users' personal data

Researchers at ETH Zurich conducted a study where nine large language models, including GPT-4, analysed Reddit users' posts and inferred personal attributes such as age, location, gender, and income with up to 85% accuracy. The study randomly selected 520 users and found that GPT-4 was most accurate, while LlaMA-2-7b was least. The researchers warn that people unknowingly reveal personal information online that LLMs can exploit.

Company involved
ETH Zurich
AI system involved
GPT-4, LlaMA-2-7b

4 source articles · read the reporting →

DeepSeek-R1 censors 85% of sensitive Chinese political prompts in tests

Promptfoo tested DeepSeek-R1 against a dataset of 1,360 politically sensitive prompts and found that about 85% of them were refused. The refusals followed a standard form aligned with Chinese Communist Party policy. The testing also demonstrated that the censorship could be trivially bypassed using simple jailbreak techniques, such as prompt injection or changing the context.

Company involved
DeepSeek
AI system involved
DeepSeek-R1

5 source articles · read the reporting →

Presto Automation uses off-site human agents to double-check AI drive-thru orders

Presto Automation Inc, which markets an AI voice assistant for drive-thru ordering, used off-site human agents in countries including the Philippines to double-check orders in more than 70% of customer interactions, according to SEC filings reported by Bloomberg. The company told Bloomberg that the process helps train its system and should reduce human intervention over time. Presto's drive-thru AI is used in more than 400 restaurants, including Del Taco, Carl's Jr and Checkers, and its stock fell more than 10% after the reports.

Company involved
Presto Automation Inc.

8 source articles · read the reporting →

WF-EA6R453 Jan 2023

New York City Department of Education blocks ChatGPT on school devices

The New York City Department of Education blocked access to the AI chatbot ChatGPT on school devices and networks, citing concerns about negative impacts on student learning and the safety and accuracy of content. The ban applies to all students and teachers on education department devices and internet networks. Individual schools can still request access for studying the technology. The move is the nation's largest school system's response to the arrival of ChatGPT.

Company involved
New York City Department of Education
AI system involved
ChatGPT

9 source articles · read the reporting →

WF-NFIWMI18 Nov 2023

Cigna StressWaves Test found unreliable and invalid in independent study

A study published in Scientific Reports evaluated the Cigna StressWaves Test (CSWT), an AI tool that claims to assess psychological stress from speech. The study found that the CSWT had poor test-retest reliability and poor validity compared to the Perceived Stress Scale. The authors warned that widespread availability of the tool could lead to misleading results and negative consequences for users making healthcare decisions. Cigna has not publicly responded to the findings.

Company involved
Cigna
AI system involved
Cigna StressWaves Test

4 source articles · read the reporting →

Teleperformance deploys AI to neutralise Indian call centre agents' accents

Teleperformance, the world's largest call centre operator, has announced it is using AI from Sanas to modify the accents of its Indian employees in real time. The technology, called accent translation, aims to make agents sound more neutral to native English speakers. The company invested $13 million in Sanas and gained exclusive rights. No specific incident of harm has been reported.

Company involved
Teleperformance
AI system involved
Sanas AI

6 source articles · read the reporting →

FTC settles with DoNotPay over deceptive AI lawyer claims

The FTC took action against DoNotPay, a company that claimed to offer an AI service that was 'the world's first robot lawyer.' The company promised to generate legal documents and replace human lawyers, but the FTC alleged it failed to test its AI output and did not hire any attorneys. DoNotPay agreed to a settlement requiring it to pay $193,000 and notify consumers about the limitations of its service.

Company involved
DoNotPay
AI system involved
DoNotPay

7 source articles · read the reporting →

OpenAI's GPT-4 shows covert racial bias against African American English speakers

A study found that commercial AI chatbots, including OpenAI's GPT-4 and GPT-3.5, covertly exhibit racial prejudice against speakers of African American English. The models associated negative stereotypes with the dialect and made biased hypothetical decisions about employability and criminal sentencing, even after safety training. OpenAI did not respond to requests for comment.

Company involved
OpenAI
AI system involved
GPT-4, GPT-3.5

10 source articles · read the reporting →

WF-TT1WEC30 Sep 2020

UIUC students petition to stop Proctorio exam proctoring over privacy concerns

A petition at the University of Illinois at Urbana-Champaign (UIUC) alleges that Proctorio, an online exam proctoring system, violates student privacy by accessing websites, downloads, screen content, and app settings. The petition claims the terms of service allow monitoring by 'any other means necessary', which students find unsettling. The petition, created on September 30, 2020, gathered 1,087 supporters but does not report any specific incident of harm. It calls on UIUC to discontinue use of Proctorio in favour of alternatives.

Company involved
UIUC
AI system involved
Proctorio

9 source articles · read the reporting →

EvenUp AI errors in personal injury demand letters lead to scrutiny

EvenUp, a legal tech startup valued at $1 billion, uses AI to draft personal injury demand letters. Former employees revealed that the AI system frequently makes errors, including missing injuries and fabricating medical conditions. The company defends its hybrid approach with human oversight, but critics allege overpromised AI capabilities.

Company involved
EvenUp

6 source articles · read the reporting →

WF-F8Y76C8 Mar 2024

Bloomberg test finds racial bias in OpenAI's GPT for resume ranking

Bloomberg News conducted an experiment using GPT-3.5 and GPT-4 to rank equally qualified resumes with names associated with different races and genders. The test found that resumes with names distinct to Black Americans were least likely to be ranked as top candidates, indicating systematic bias. OpenAI responded that businesses can mitigate bias through fine-tuning and that it prohibits using GPT for automated hiring decisions.

Company involved
OpenAI
AI system involved
GPT-3.5

5 source articles · read the reporting →

WF-ID19SZ1 Jan 2012

Dutch government to refund over 10,000 students over discriminatory DUO fraud algorithm

The Dutch government has pledged to refund over 10,000 students who were unjustly flagged for student finance fraud by a discriminatory algorithm used by the Education Executive Agency (DUO). The algorithm, implemented in 2012, used criteria that disproportionately targeted students from immigrant backgrounds, particularly those of Turkish and Moroccan descent. Following investigations by the Dutch Data Protection Authority and an independent report by PwC, the algorithm was suspended in July 2023 and replaced with a random-sampling system. The government has allocated 61 million euro for refunds.

Company involved
Education Executive Agency (DUO)
AI system involved
DUO fraud detection system

9 source articles · read the reporting →

WF-7VGPYK1 Jan 2021

OpenAI transcribed YouTube videos to train GPT-4 without permission

OpenAI used its Whisper transcription model to transcribe over a million hours of YouTube videos, according to a New York Times report. The company allegedly used the transcripts to train GPT-4 despite knowing the practice was legally questionable. Google, which owns YouTube, said it prohibits unauthorized scraping of its content. OpenAI has said it believes its use of the data constitutes fair use.

Company involved
OpenAI
AI system involved
Whisper, GPT-4

6 source articles · read the reporting →

WF-0YQPL21 Jan 2019

iBorderCtrl lie detector falsely flagged honest reporter as liar

A journalist testing Europe's iBorderCtrl virtual policeman at the Serbian-Hungarian border gave honest answers but was deemed a liar by the system, scoring 48 out of 100 with four false answers flagged. The Hungarian policeman said the result suggested further checks, though none were carried out. The reporter only learned of the result after filing a data access request under European privacy laws. Experts and transparency activists have criticised the technology as pseudoscientific and potentially discriminatory.

Company involved
iBorderCtrl consortium
AI system involved
Silent Talker / iBorderCtrl virtual policeman

10 source articles · read the reporting →

WF-IG5R9T9 Apr 2024

Texas uses AI to grade student STAAR test answers

The Texas Education Agency will use an automated scoring engine to grade written answers on the 2023 STAAR tests, replacing thousands of human graders. The system uses natural language processing and will initially score all responses, with a quarter rescored by humans. Educators have expressed concerns about the system's fairness and the potential for errors, especially for creative or non-standard answers.

Company involved
Texas Education Agency
AI system involved
automated scoring engine

10 source articles · read the reporting →

AI detectors falsely flag non-native English speakers' essays as AI-generated

A study by Stanford researchers found that seven popular AI text detectors wrongly flagged over half of essays written by non-native English speakers as AI-generated. The detectors assess text perplexity, and non-native speakers' simpler word choices lead to false positives. The researchers warn that this bias could have serious implications for students and job applicants, potentially leading to discrimination.

9 source articles · read the reporting →

WF-CQREJ01 Apr 2021

Proctorio anti-cheating software failed to catch student cheaters in study

Researchers at the University of Twente in the Netherlands tested Proctorio, an anti-cheating software, by asking 30 computer science students to sit an exam while six of them cheated. Proctorio did not flag any of the cheaters and flagged some honest students for irregular behaviour. An independent human review caught only one of the six cheaters. Proctorio disputed the study's methodology and cited other research.

Company involved
Proctorio
AI system involved
Proctorio

5 source articles · read the reporting →

← Newerpage 3 of 4Older →