The record

Where automated decisions went wrong

Incidents gathered from public reporting around the world. Each one links to the articles it came from. None of it is a finding that anyone broke the law.

Reports people file about their own experience are not shown here and never will be without their agreement. Tell us what happened to you.

Clear

92 incidents closest to “Turnitin Similarity” · matched on meaning · public reporting

WF-961A0L1 Dec 2018

Educational Testing Service's E-rater algorithm biases essay scores against minority students

The Educational Testing Service's E-rater algorithm, used to grade essays on the GRE and other standardized tests, has been found to systematically give higher scores to students from mainland China and lower scores to African American students compared to human graders. The bias stems from the algorithm's reliance on surface-level metrics like vocabulary and sentence length, which disadvantage certain groups. Despite studies dating back to 1999, the bias persists, and in many states, only a small percentage of essays are reviewed by humans.

Company involved
Educational Testing Service
AI system involved
E-rater

10 source articles · read the reporting →

WF-AYQCBO24 Oct 2024

Google, Microsoft, and Perplexity AI search results promote racist IQ data

AI-powered search engines from Google, Microsoft, and Perplexity have been surfacing debunked research promoting race science, including false IQ scores for countries. The systems pulled data from a dataset by Richard Lynn, a known proponent of scientific racism. Google removed the offending Overviews after being contacted by WIRED, but the data still appears in featured snippets and other AI tools.

Company involved
Google, Microsoft, and Perplexity
AI system involved
AI Overviews, Copilot, Perplexity

7 source articles · read the reporting →

WF-TNB7Q615 Oct 2024

NSW Education Standards Authority used AI-generated image in HSC English exam without disclosure

The NSW Education Standards Authority (NESA) used an AI-generated image as a stimulus in the 2024 HSC English exam without disclosing its origin. The image, created by Florian Schroeder using OpenAI's ChatGPT and Dall-E 2, was published on Medium in July 2023. Students suspected AI use due to irregularities in the image, and NESA initially declined to confirm. After the Sydney Morning Herald confirmed the image was AI-generated, NESA stated that students would be marked on their response to the question, not the image's origin.

Company involved
NSW Education Standards Authority
AI system involved
ChatGPT and Dall-E 2

6 source articles · read the reporting →

WF-W8MY6X1 Sep 2020

Middle schooler beats Edgenuity grading algorithm to get perfect score

A seventh-grade student in the Los Angeles Unified School District received a failing grade on a history assignment graded by Edgenuity's automated scoring algorithm. With help from his mother, a history professor, he reverse-engineered the algorithm by writing a paragraph with relevant keywords and a jumble of words, earning a perfect score. The incident highlights concerns about the accuracy and fairness of automated grading systems in education.

Company involved
Los Angeles Unified School District
AI system involved
Edgenuity

10 source articles · read the reporting →

WF-XH2W9I1 Sep 2020

Proctortrack data breach exposed student data from online proctoring

Proctortrack, an online proctoring service used by universities, suffered a data breach in September 2020 when its source code was leaked online. An analysis by Consumer Reports found that the code contained hard-coded passwords and exposed the names and email addresses of over 150 students. The company acknowledged the leak but said no harm resulted. Students had been required to use the software, which performed facial recognition and recorded video during exams.

Company involved
Proctortrack
AI system involved
Proctortrack

10 source articles · read the reporting →

WF-Q99LTG1 Aug 2020

College student uses GPT-3 to create fake blog that reaches #1 on Hacker News

Liam Porr, a college student at UC Berkeley, used OpenAI's GPT-3 language model to generate a fake blog under a fake name. One of the posts reached the number-one spot on Hacker News, fooling tens of thousands of readers. Porr later confessed and retired the blog after two weeks. The experiment demonstrated the ease of creating convincing AI-generated content.

AI system involved
GPT-3

10 source articles · read the reporting →

WF-OH4XSX22 Mar 2023

ChatGPT-4 generates false accusations and quotes about law professors

OpenAI's ChatGPT-4 generated false accusations and fabricated quotes about law professors when asked about scandals and crimes involving law professors. The system produced entirely made-up newspaper citations and quotes, falsely alleging misconduct such as harassment and tax fraud. The article reports that these hallucinations could mislead users who trust the generated quotes. The author corrected an earlier error attributing similar results to ChatGPT-4 instead of ChatGPT-3.5, but confirmed that ChatGPT-4 also produces such false outputs.

Company involved
OpenAI
AI system involved
ChatGPT-4

10 source articles · read the reporting →

WF-2N0W571 Mar 2023

NewsGuard finds ChatGPT-4 generates all 100 false narratives in test

In March 2023, NewsGuard tested ChatGPT-4 by prompting it with 100 false narratives from its database. The chatbot generated all 100 false narratives, often without disclaimers, and produced more persuasive misinformation than its predecessor ChatGPT-3.5. NewsGuard reported that OpenAI had not fixed the flaw before releasing the tool, and the company did not respond to requests for comment.

Company involved
OpenAI
AI system involved
ChatGPT-4

3 source articles · read the reporting →

WF-701YEA11 Mar 2021

UW-Madison disables Honorlock after skin tone recognition failure

The University of Wisconsin-Madison disabled the exam pause feature of its Honorlock anti-cheating software in March 2021 after three students complained that the software failed to recognize their darker skin tones and paused their exams. The software, used since online classes began, automatically pauses exams when it cannot detect facial features. Honorlock denied the issue was related to skin tone, attributing it to students looking away from their webcams. The university responded by disabling the feature.

Company involved
University of Wisconsin-Madison
AI system involved
Honorlock

10 source articles · read the reporting →

Dartmouth Medical School charges 17 students with cheating based on tracking system.

Dartmouth College's Geisel School of Medicine charged 17 students with cheating on remote exams during the pandemic. The charges were based on secret tracking of student activity on the learning management system. Critics say the system is prone to errors and inappropriate for monitoring students.

Company involved
Dartmouth College Geisel School of Medicine

10 source articles · read the reporting →

Georgetown researchers use GPT-3 to generate convincing misinformation

Researchers at Georgetown University's Center for Security and Emerging Technology trained OpenAI's GPT-3 to generate convincing misinformation. In tests, users exposed to AI-generated messages opposing sanctions on China doubled their opposition to sanctions. The research demonstrates the potential for AI to spread disinformation.

Company involved
Georgetown University
AI system involved
GPT-3

10 source articles · read the reporting →

WF-HY5JT527 Nov 2024

Tow Center finds ChatGPT Search misattributes publisher content

The Tow Center for Digital Journalism tested ChatGPT Search with 200 block quotes from 20 publishers and found 153 partially or fully incorrect citations. The chatbot often conjured responses when it could not access content, sometimes citing plagiarized or syndicated versions. OpenAI responded that the study was atypical and that it supports publishers with clear links and attribution.

Company involved
OpenAI
AI system involved
ChatGPT Search

6 source articles · read the reporting →

WF-YWUJN81 Oct 2024

Company fires HR team after ATS auto-rejects manager's CV due to filtering error

A company's applicant tracking system (ATS) auto-rejected qualified candidates' resumes for three months because it was filtering for the outdated framework AngularJS instead of the required Angular framework. The manager discovered the flaw by submitting his own CV under a pseudonym and found it was rejected within seconds. After the manager reported the issue to upper management, the company investigated and dismissed half of its HR team. No legal action or regulatory involvement is reported.

4 source articles · read the reporting →

WF-3UCXWE1 Jul 2023

Study finds Midjourney, DALL-E 2, Stable Diffusion accept over 85% of fake news prompts

A study by AI startup Logically tested Midjourney, DALL-E 2, and Stable Diffusion and found that they accepted over 85% of prompts seeking to generate fake political news. The systems generated images of ballot stuffing, small boat arrivals, and explosions. Logically warned that the lack of moderation could pose threats to upcoming elections. Stability AI responded by stating its ethical use license and measures to prevent misuse.

Company involved
Midjourney, OpenAI, Stability AI
AI system involved
Midjourney, DALL-E 2, Stable Diffusion

8 source articles · read the reporting →

WF-O8P3831 May 2023

GlobalVillageSpace.com used AI to rewrite New York Times articles without credit

NewsGuard identified 37 websites using AI chatbots to rewrite articles from mainstream news outlets without credit. One example is GlobalVillageSpace.com, which appeared to use AI to rewrite a New York Times article about NFL tight end Darren Waller. The site published an AI error message indicating the article was rewritten. After NewsGuard contacted the site, it removed the article but did not respond to inquiries.

Company involved
GlobalVillageSpace.com

6 source articles · read the reporting →

DeepSeek-R1 censors 85% of sensitive Chinese political prompts in tests

Promptfoo tested DeepSeek-R1 against a dataset of 1,360 politically sensitive prompts and found that about 85% of them were refused. The refusals followed a standard form aligned with Chinese Communist Party policy. The testing also demonstrated that the censorship could be trivially bypassed using simple jailbreak techniques, such as prompt injection or changing the context.

Company involved
DeepSeek
AI system involved
DeepSeek-R1

5 source articles · read the reporting →

Ubisoft faces backlash after announcing Ghostwriter AI writing tool

Ubisoft has announced Ubisoft Ghostwriter, an AI tool that generates first drafts of non-player character dialogue. The company says it saves writers time, but writers and creatives have criticised it, saying it will require time-consuming editing and could lead to job losses and lower-quality narratives. Ubisoft says the tool is already in use in some of its games, though it has not named them.

Company involved
Ubisoft
AI system involved
Ubisoft Ghostwriter

10 source articles · read the reporting →

WF-51NEWF17 Feb 2024

UIUC researchers weaponize GPT-4 to autonomously hack websites

Researchers at the University of Illinois Urbana-Champaign demonstrated that LLM-powered agents, particularly OpenAI's GPT-4, can autonomously hack vulnerable websites. In sandboxed tests, GPT-4 achieved a 73.3% success rate across five attempts on 15 vulnerabilities, while open-source models failed. The researchers used the OpenAI Assistants API, LangChain, and Playwright to enable the agents to interact with websites. The study highlights the potential for AI agents to be used in cyberattacks, with cost estimates suggesting they could be cheaper than human penetration testers.

Company involved
University of Illinois Urbana-Champaign
AI system involved
GPT-4

4 source articles · read the reporting →

WF-F8Y76C8 Mar 2024

Bloomberg test finds racial bias in OpenAI's GPT for resume ranking

Bloomberg News conducted an experiment using GPT-3.5 and GPT-4 to rank equally qualified resumes with names associated with different races and genders. The test found that resumes with names distinct to Black Americans were least likely to be ranked as top candidates, indicating systematic bias. OpenAI responded that businesses can mitigate bias through fine-tuning and that it prohibits using GPT for automated hiring decisions.

Company involved
OpenAI
AI system involved
GPT-3.5

5 source articles · read the reporting →

Academic journals publish papers with AI-generated text from ChatGPT

Scientific journals have published papers containing text that appears to have been generated by AI tools like ChatGPT. A search for the phrase 'As of my last knowledge update' on Google Scholar returned 115 results, indicating that researchers or authors used ChatGPT to write parts of their papers. The phrase is characteristic of ChatGPT's responses and corresponds to its knowledge update dates. The incident highlights the pervasive use of AI in academic publishing and raises concerns about the integrity of peer-reviewed literature.

Company involved
Academic journals
AI system involved
ChatGPT

10 source articles · read the reporting →

WF-UENLG626 Mar 2024

ChatGPT use linked to memory loss and procrastination in students

A study published in the International Journal of Educational Technology in Higher Education surveyed hundreds of university students in Pakistan and found that those who relied more on ChatGPT reported increased procrastination, memory loss, and lower GPAs. The researchers attribute this to the chatbot making schoolwork too easy, reducing students' cognitive effort. The study's lead author warned of a "dark side" to excessive generative AI usage.

Company involved
National University of Computer and Emerging Sciences
AI system involved
ChatGPT

4 source articles · read the reporting →

Study finds LLMs used in up to 16.9% of AI conference peer reviews

According to a new paper on arXiv, researchers have begun using generative AI services to help write peer reviews of machine learning papers submitted to leading AI conferences. The study analysed reviews from ICLR 2024, NeurIPS 2023, CoRL 2023 and EMNLP 2023 and estimated that between 6.5% and 16.9% of review text may have been substantially modified by large language models. The authors argue that this risks depriving authors of diverse expert feedback and may skew reviews towards AI model biases. They have called for greater transparency about the use of LLMs in peer review.

9 source articles · read the reporting →

WF-G4FI1630 Apr 2025

Anthropic ordered to respond over alleged AI-hallucinated citation in court filing

A US federal magistrate judge has ordered Anthropic to respond to music publishers' claim that a court filing by Anthropic data scientist Olivia Chen cited a fictitious academic article that may have been generated by Anthropic's AI tool Claude. The publishers' lawyer said he had confirmed with the named author and journal that the article did not exist. Anthropic's lawyer disputed this, saying it was a mis-citation rather than an AI hallucination. The filing was made in an ongoing copyright case brought by Universal Music Group, Concord, and ABKCO against Anthropic over the use of song lyrics to train Claude.

Company involved
Anthropic
AI system involved
Claude

5 source articles · read the reporting →

WF-IG5R9T9 Apr 2024

Texas uses AI to grade student STAAR test answers

The Texas Education Agency will use an automated scoring engine to grade written answers on the 2023 STAAR tests, replacing thousands of human graders. The system uses natural language processing and will initially score all responses, with a quarter rescored by humans. Educators have expressed concerns about the system's fairness and the potential for errors, especially for creative or non-standard answers.

Company involved
Texas Education Agency
AI system involved
automated scoring engine

10 source articles · read the reporting →

← Newerpage 3 of 4Older →