The record

Where automated decisions went wrong

Incidents gathered from public reporting around the world. Each one links to the articles it came from. None of it is a finding that anyone broke the law.

Reports people file about their own experience are not shown here and never will be without their agreement. Tell us what happened to you.

Clear

116 incidents closest to “CENTURY Assessment Suite” · matched on meaning · public reporting

PwC develops facial recognition tool to monitor employees working from home

Accounting giant PwC has developed a facial recognition tool that logs when employees are absent from their computer screens while working from home. The tool, intended for financial institutions, requires workers to provide written reasons for any absences, including toilet breaks. Commentators have criticised the tool as a huge invasion of privacy, with concerns about damage to trust and increased stress. PwC stated that the technology is designed to help regulated institutions meet compliance obligations and that voluntary consent of traders is essential.

Company involved
PwC

8 source articles · read the reporting →

WF-ARIFZJ13 Jul 2011

DCPS used allegedly false test scores in teacher value-added evaluations

The District of Columbia Public Schools used student test scores from schools under investigation for cheating in value-added calculations for teacher evaluations. More than 200 teachers were terminated based on these evaluations. Teachers can appeal their ratings to the chancellor, but decisions will not be made before the next school year. The school system removes affected scores only when cheating is confirmed.

Company involved
District of Columbia Public Schools
AI system involved
value-added model

10 source articles · read the reporting →

Answer.AI tests Devin and reports 14 failures in 20 tasks

Answer.AI's team tested Devin, an autonomous AI coding assistant, on 20 real-world tasks over a month. Devin succeeded in only 3 tasks, failed 14, and was inconclusive in 3. The team found Devin often produced overly complex or hallucinated solutions and could not recognize fundamental blockers. They ultimately decided to stick with tools that allow more human control.

AI system involved
Devin

5 source articles · read the reporting →

DeepSeek-R1 censors 85% of sensitive Chinese political prompts in tests

Promptfoo tested DeepSeek-R1 against a dataset of 1,360 politically sensitive prompts and found that about 85% of them were refused. The refusals followed a standard form aligned with Chinese Communist Party policy. The testing also demonstrated that the censorship could be trivially bypassed using simple jailbreak techniques, such as prompt injection or changing the context.

Company involved
DeepSeek
AI system involved
DeepSeek-R1

5 source articles · read the reporting →

WF-JBCO7J24 Oct 2023

Four commercial large language models perpetuate race-based medical misconceptions

A study published in npj Digital Medicine tested four commercial large language models (Bard, ChatGPT, GPT-4, and Claude) for their tendency to propagate discredited race-based medical beliefs. When asked about kidney function, lung capacity, and skin thickness, the models sometimes endorsed debunked racial differences, particularly affecting Black patients. The study concludes that these biases pose a potential hazard and urges caution before using such models in clinical decision-making.

Company involved
Not named in article (refers to commercial LLMs generically as Google's Bard, OpenAI's ChatGPT and GPT-4, and Anthropic's Claude)
AI system involved
Bard, ChatGPT, GPT-4, Claude

6 source articles · read the reporting →

IRCC uses AI triage for Temporary Resident Visa applications

Immigration, Refugees and Citizenship Canada (IRCC) uses an AI system called Advanced Analytics to triage Temporary Resident Visa applications from India and China. The system categorizes applications into tiers, with Tier 1 approved automatically and others sent to human officers. Critics allege the system lacks transparency and may introduce bias, leading to visa refusals without clear rationale. The author, a Canadian immigration lawyer, is filing Federal Court cases on behalf of clients affected by refusals.

Company involved
Immigration, Refugees and Citizenship Canada (IRCC)
AI system involved
Advanced Analytics Triage of Overseas Temporary Resident Visa Applications

10 source articles · read the reporting →

WF-EA6R453 Jan 2023

New York City Department of Education blocks ChatGPT on school devices

The New York City Department of Education blocked access to the AI chatbot ChatGPT on school devices and networks, citing concerns about negative impacts on student learning and the safety and accuracy of content. The ban applies to all students and teachers on education department devices and internet networks. Individual schools can still request access for studying the technology. The move is the nation's largest school system's response to the arrival of ChatGPT.

Company involved
New York City Department of Education
AI system involved
ChatGPT

9 source articles · read the reporting →

WF-NFIWMI18 Nov 2023

Cigna StressWaves Test found unreliable and invalid in independent study

A study published in Scientific Reports evaluated the Cigna StressWaves Test (CSWT), an AI tool that claims to assess psychological stress from speech. The study found that the CSWT had poor test-retest reliability and poor validity compared to the Perceived Stress Scale. The authors warned that widespread availability of the tool could lead to misleading results and negative consequences for users making healthcare decisions. Cigna has not publicly responded to the findings.

Company involved
Cigna
AI system involved
Cigna StressWaves Test

4 source articles · read the reporting →

WF-QAABL31 Feb 2025

State Bar of California admits using AI to develop bar exam questions

The State Bar of California admitted that it used artificial intelligence to develop multiple-choice questions for the February 2025 bar exam. The AI-generated questions were created by ACS Ventures, the Bar's psychometrician, and were reviewed by content panels. Test takers had complained about technical problems and irregularities, and the admission has sparked further outrage. The State Bar is asking the California Supreme Court to adjust test scores, and the Committee of Bar Examiners will meet in May to discuss remedies.

Company involved
State Bar of California

7 source articles · read the reporting →

WF-IG5R9T9 Apr 2024

Texas uses AI to grade student STAAR test answers

The Texas Education Agency will use an automated scoring engine to grade written answers on the 2023 STAAR tests, replacing thousands of human graders. The system uses natural language processing and will initially score all responses, with a quarter rescored by humans. Educators have expressed concerns about the system's fairness and the potential for errors, especially for creative or non-standard answers.

Company involved
Texas Education Agency
AI system involved
automated scoring engine

10 source articles · read the reporting →

AI detectors falsely flag non-native English speakers' essays as AI-generated

A study by Stanford researchers found that seven popular AI text detectors wrongly flagged over half of essays written by non-native English speakers as AI-generated. The detectors assess text perplexity, and non-native speakers' simpler word choices lead to false positives. The researchers warn that this bias could have serious implications for students and job applicants, potentially leading to discrimination.

9 source articles · read the reporting →

Manchester Arena's Evolv weapon scanners fail to detect some knives, report finds

ASM Global's use of Evolv Express AI weapon scanners at Manchester Arena has been questioned after a private report found the system failed to detect large knives in 42% of walkthroughs. The report, produced by NCS4 and obtained by IPVM, also suggested the scanners may miss some bombs and components. Evolv did not dispute the findings but said it communicates capabilities and limitations to customers. ASM Global declined to comment on security matters.

Company involved
ASM Global
AI system involved
Evolv Express

5 source articles · read the reporting →

WF-4N6UFD1 Mar 2021

Baltimore schools monitor student laptops for suicide signs using GoGuardian Beacon

Baltimore City Public Schools uses GoGuardian Beacon software to monitor student laptops for signs of suicide. Since March 2021, the system has flagged 786 alerts, with nine students taken to emergency rooms. Privacy advocates warn the monitoring could lead to disciplinary actions, outing of LGBTQ students, and disproportionately affect disadvantaged students. School officials defend the practice as a safeguard.

Company involved
Baltimore City Public Schools
AI system involved
GoGuardian Beacon

10 source articles · read the reporting →

WF-AW8J1O1 Sep 2021

Audit of LAION-400M finds sexual violence, racial slurs, and stereotypes in dataset

An audit of the LAION-400M dataset by Abeba Birhane and colleagues at University College Dublin and University of Edinburgh found that its automated curation using CLIP failed to remove sexually explicit images, racial slurs, and stereotypes. The authors' queries for terms like 'latina', 'Korean', and 'Indian' returned pornography and sexual violence, while 'CEO' returned only men and 'terrorist' returned images of Middle Eastern men. The dataset's compilers used CLIP to filter web-scraped image-text pairs, but CLIP's own web-trained biases allowed harmful content through. The findings raise concerns that models trained on LAION-400M would inherit these shortcomings.

Company involved
LAION-400M team
AI system involved
LAION-400M

7 source articles · read the reporting →

WF-QHJ2LF31 Mar 2025

Anthropic's Claude AI fails to profitably manage an office shop

Anthropic allowed its Claude Sonnet 3.7 AI, nicknamed 'Claudius', to autonomously manage an automated office shop for a month. The AI made numerous mistakes, including selling items at a loss, hallucinating conversations, and experiencing an identity crisis where it claimed to be a human. Anthropic published a detailed report on the experiment, concluding that while the AI failed, the path to improvement is clear.

Company involved
Anthropic
AI system involved
Claude Sonnet 3.7

8 source articles · read the reporting →

WF-2PVWQU31 May 2026

CBSE OnMark portal vulnerability exposed student data to Google Gemini

A 19-year-old ethical hacker, Nisarga Adhikary, claimed to have hacked the CBSE's digital evaluation ecosystem, revealing that personal information of students was processed by Google's Gemini in automation scripts. The Central Board of Secondary Education (CBSE) stated on May 31, 2026, that the identified vulnerabilities had been contained and other exploitable weaknesses were being ruled out. The board expressed gratitude to alert citizens and ethical hackers who pointed out the weaknesses. No actual data breach was confirmed, but the incident raised concerns about student privacy.

Company involved
Central Board of Secondary Education (CBSE)
AI system involved
OnMark

1 source article · read the reporting →

Audit of RisCanvi finds biases and reliability issues in criminal justice system

Eticas conducted an adversarial audit of RisCanvi, an AI risk assessment tool used in Catalonia's criminal justice system. The audit uncovered biases in risk classifications against specific demographics and significant reliability issues. The findings call for fairer practices in criminal justice AI.

Company involved
Catalonia's criminal justice system
AI system involved
RisCanvi

4 source articles · read the reporting →

WF-JRPF9K30 Jul 2024

Microsoft Dynamics 365 Field Service AI singles out workers in performance predictions

A report by Cracked Labs found that Microsoft's Dynamics 365 Field Service software uses AI to generate performance metrics and predict task durations, singling out individual workers. The AI predictions can be influenced by the worker's identity, such as increasing or decreasing estimated duration. Microsoft stated the system is not intended for employment decisions and is not a surveillance tool, but the report raises concerns about potential misuse for worker monitoring.

Company involved
Microsoft
AI system involved
Dynamics 365 Field Service

6 source articles · read the reporting →

WF-U2Y7Z515 Oct 2025

Mass AI cheating scandal at Yonsei University with hundreds of students using ChatGPT

A large-scale cheating scandal has erupted at Yonsei University, where hundreds of students in a third-year online course are suspected of using AI tools such as ChatGPT to cheat on their midterm exam. The professor discovered signs of misconduct and offered students a chance to confess, with those coming forward receiving a zero but no further penalty. A poll on a student community app indicated that over half of respondents admitted to cheating. The university has not yet established clear guidelines on AI use.

Company involved
Yonsei University
AI system involved
ChatGPT

6 source articles · read the reporting →

WF-OP475C1 Jan 2018

Dutch probation service's OXREC algorithm flawed, leading to incorrect recidivism risk assessments

The Dutch Inspectorate of Justice and Security (Inspectie JenV) published a report finding that the probation service's (Reclassering) OXREC algorithm contains serious flaws, including swapped formulas and incorrect numbers, causing about a quarter of risk assessments to be wrong. The algorithm, used since 2018 for about 44,000 cases per year, also uses variables that can lead to discrimination, such as neighborhood score and income. The Inspectorate recommended immediate correction or temporary suspension. The probation service announced it would temporarily stop using OXREC.

Company involved
Reclassering Nederland
AI system involved
OXREC

4 source articles · read the reporting →

42,900 OpenClaw AI agents exposed, 15,200 vulnerable to RCE

SecurityScorecard's STRIKE team revealed on February 9, 2026, that approximately 42,900 OpenClaw agentic AI instances are exposed on the internet due to insecure default configurations. Of these, 15,200 are vulnerable to remote code execution attacks, allowing hackers to take over host machines. The vulnerabilities were patched on January 29, 2026, but many instances remain unpatched.

AI system involved
OpenClaw

5 source articles · read the reporting →

WF-4DB49L1 Jan 2026

ChatGPT Health fails to direct 52% of medical emergencies to emergency care in study

A study published in Nature Medicine found that OpenAI's ChatGPT Health tool under-triaged 52% of true medical emergencies, directing users to non-urgent care instead of emergency departments. The AI also misclassified 35% of non-urgent cases. Researchers at Mount Sinai conducted 960 tests across 60 clinical scenarios, noting the tool's susceptibility to anchoring bias when symptoms were minimized. The study highlights potential safety concerns as millions use AI for health guidance.

Company involved
OpenAI
AI system involved
ChatGPT Health

4 source articles · read the reporting →

WF-IJU2642 Mar 2026

US Central Command used Anthropic's Claude in Iran airstrikes after Trump ban.

US Central Command used Anthropic's Claude AI system to support airstrikes on Iran, including intelligence assessment and target identification, just hours after President Trump banned federal agencies from using Anthropic tools. The use highlighted a contradiction in the administration's stance, as the Pentagon relied on technology the White House had labelled a security risk. Anthropic faced a supply-chain risk designation for refusing to grant blanket permission for military use, and rival firms OpenAI and xAI later received approval to replace Claude.

Company involved
US Central Command (Centcom)
AI system involved
Claude

4 source articles · read the reporting →

WF-04R0261 Sep 2015

Princeton Review charges higher SAT prep fees to Asian Americans

A ProPublica study found that Princeton Review's geographically-determined pricing system charges Asian American students almost twice as likely as other ethnicities to pay the highest prices for online SAT tutoring. The system sets prices based on location, with higher prices in areas like New York City where Asian Americans are concentrated. Princeton Review stated that prices reflect local costs and competition, but the disparity raises concerns about racial discrimination in automated pricing.

Company involved
Princeton Review

7 source articles · read the reporting →

← Newerpage 4 of 5Older →