The record

Where automated decisions went wrong

Incidents gathered from public reporting around the world. Each one links to the articles it came from. None of it is a finding that anyone broke the law.

Reports people file about their own experience are not shown here and never will be without their agreement. Tell us what happened to you.

Clear

92 incidents closest to “Respondus 4” · matched on meaning · public reporting

WF-II25VQ28 Jul 2020

ExamSoft failure during Michigan Bar Exam

On July 28, 2020, ExamSoft's software platform failed during the Michigan Bar Exam, preventing some test takers from completing the exam. ExamSoft posted a statement on X (formerly Twitter) with the hashtag #MichiganBarExam, but critics noted the company deleted and reposted the statement multiple times, allegedly to hide negative comments. The incident affected a group of examinees and caused disruption to their ability to take the exam.

Company involved
ExamSoft
AI system involved
ExamSoft platform

10 source articles · read the reporting →

WF-SD7T9128 Sep 2020

Student flagged by ProctorU for reading aloud during exam

A college student, Dana Jo, was flagged by ProctorU test proctoring software for talking during an exam, which she says was reading a question aloud. Her professor initially gave her a zero and placed an academic infraction on her record, jeopardizing her scholarships. After reviewing a video recording, the professor apologized, reinstated her grade, and removed the infraction. ProctorU's CEO stated that the incident highlights the importance of video recordings for review.

Company involved
University (not named)
AI system involved
ProctorU

5 source articles · read the reporting →

Clearview AI tested facial recognition surveillance cameras with UFT and Rudin

Clearview AI, the facial recognition company that scraped billions of photos from social media, developed a surveillance camera system under the name Insight Camera. The system was tested by the United Federation of Teachers and Rudin Management in New York City. The UFT used it to identify individuals who had made threats and prevent them from entering its offices. Clearview did not respond to requests for comment.

Company involved
Clearview AI
AI system involved
Insight Camera

9 source articles · read the reporting →

WF-OH4XSX22 Mar 2023

ChatGPT-4 generates false accusations and quotes about law professors

OpenAI's ChatGPT-4 generated false accusations and fabricated quotes about law professors when asked about scandals and crimes involving law professors. The system produced entirely made-up newspaper citations and quotes, falsely alleging misconduct such as harassment and tax fraud. The article reports that these hallucinations could mislead users who trust the generated quotes. The author corrected an earlier error attributing similar results to ChatGPT-4 instead of ChatGPT-3.5, but confirmed that ChatGPT-4 also produces such false outputs.

Company involved
OpenAI
AI system involved
ChatGPT-4

10 source articles · read the reporting →

WF-PMLNUT17 Sep 2020

California bar exam facial recognition system fails to verify Arab-American student's identity

Ahmed Alamri, an Arab-American law student, was unable to register for the practice California bar exam because ExamSoft's facial recognition system repeatedly failed to recognize his face, citing poor lighting. Alamri attempted to verify his identity over 75 times in different rooms and lighting conditions without success. He and other students filed an emergency petition with the California Supreme Court, alleging that the system discriminates against people of color. The incident highlights concerns about bias in facial recognition technology used for exam proctoring.

Company involved
California State Bar
AI system involved
ExamSoft

10 source articles · read the reporting →

WF-2N0W571 Mar 2023

NewsGuard finds ChatGPT-4 generates all 100 false narratives in test

In March 2023, NewsGuard tested ChatGPT-4 by prompting it with 100 false narratives from its database. The chatbot generated all 100 false narratives, often without disclaimers, and produced more persuasive misinformation than its predecessor ChatGPT-3.5. NewsGuard reported that OpenAI had not fixed the flaw before releasing the tool, and the company did not respond to requests for comment.

Company involved
OpenAI
AI system involved
ChatGPT-4

3 source articles · read the reporting →

Cleveland State University's room scan requirement ruled unconstitutional

A federal judge ruled that Cleveland State University's requirement for a student to undergo a 360-degree room scan before an online exam was an unreasonable search under the Fourth Amendment. The student, enrolled at the public university, was told shortly before the exam that he would need to scan his private space. The court found that the university's justifications did not outweigh the privacy protections of the home. No final judgment or injunction has been issued yet.

Company involved
Cleveland State University

10 source articles · read the reporting →

WF-701YEA11 Mar 2021

UW-Madison disables Honorlock after skin tone recognition failure

The University of Wisconsin-Madison disabled the exam pause feature of its Honorlock anti-cheating software in March 2021 after three students complained that the software failed to recognize their darker skin tones and paused their exams. The software, used since online classes began, automatically pauses exams when it cannot detect facial features. Honorlock denied the issue was related to skin tone, attributing it to students looking away from their webcams. The university responded by disabling the feature.

Company involved
University of Wisconsin-Madison
AI system involved
Honorlock

10 source articles · read the reporting →

WF-ZQ3YFL1 Dec 2020

Retorio AI personality test swayed by candidate appearance in BR experiment

Bayerischer Rundfunk journalists conducted experiments with Retorio's AI video interview analysis tool. The AI, which assesses personality traits from short videos, produced different scores when the same actress changed her appearance (glasses, headscarf, wig) or the video background and lighting were altered. The start-up Retorio acknowledged that the AI considers external image, similar to a human interviewer. Experts warned that such software could perpetuate stereotypes and unfairly affect job candidates.

AI system involved
Retorio AI

10 source articles · read the reporting →

Dartmouth Medical School charges 17 students with cheating based on tracking system.

Dartmouth College's Geisel School of Medicine charged 17 students with cheating on remote exams during the pandemic. The charges were based on secret tracking of student activity on the learning management system. Critics say the system is prone to errors and inappropriate for monitoring students.

Company involved
Dartmouth College Geisel School of Medicine

10 source articles · read the reporting →

WF-LKR1LN18 Aug 2026

Joann LeDoux v. Outliers, Inc. (2) (W.D. Washington): AI-hallucinated content in court filing, Expert Brief excluded/struck

The AI-generated hallucinated citations in an expert report led to the exclusion of the expert and dismissal of the plaintiff's case.

1 source article · read the reporting →

WF-XO6C6612 Jul 2021

Userviz machine-learning aimbot shut down after Activision request

A developer known as User101 shut down the Userviz cheat after Activision requested that he stop developing it. The software used computer vision and machine learning to automate aiming in games such as Call of Duty: Warzone, and was marketed as undetectable. It was never published, and the developer said his intention was not to do anything illegal.

Company involved
User101
AI system involved
Userviz

7 source articles · read the reporting →

WF-4RKE3V1 May 2016

Danish students publish OkCupid user data, DPA investigates

Two Danish students scraped and published data on 70,000 OkCupid users, including sensitive sexual preferences and religious views, without anonymisation. The dataset was removed after OkCupid filed a copyright notice. The Danish Data Protection Authority (Datatilsynet) has launched an investigation into the incident.

10 source articles · read the reporting →

WF-YWUJN81 Oct 2024

Company fires HR team after ATS auto-rejects manager's CV due to filtering error

A company's applicant tracking system (ATS) auto-rejected qualified candidates' resumes for three months because it was filtering for the outdated framework AngularJS instead of the required Angular framework. The manager discovered the flaw by submitting his own CV under a pseudonym and found it was rejected within seconds. After the manager reported the issue to upper management, the company investigated and dismissed half of its HR team. No legal action or regulatory involvement is reported.

4 source articles · read the reporting →

Answer.AI tests Devin and reports 14 failures in 20 tasks

Answer.AI's team tested Devin, an autonomous AI coding assistant, on 20 real-world tasks over a month. Devin succeeded in only 3 tasks, failed 14, and was inconclusive in 3. The team found Devin often produced overly complex or hallucinated solutions and could not recognize fundamental blockers. They ultimately decided to stick with tools that allow more human control.

AI system involved
Devin

5 source articles · read the reporting →

Purdue study finds ChatGPT wrong over half the time on software questions

A study by Purdue University found that ChatGPT provided incorrect answers to over half of 517 software development questions from Stack Overflow. Despite the errors, 34% of users preferred ChatGPT's answers over human responses. The study warns that relying on ChatGPT for coding could jeopardize programmers' professional reputations.

Company involved
OpenAI
AI system involved
ChatGPT

9 source articles · read the reporting →

WF-G1A5LS1 Nov 2023

ETH Zurich study shows LLMs can infer Reddit users' personal data

Researchers at ETH Zurich conducted a study where nine large language models, including GPT-4, analysed Reddit users' posts and inferred personal attributes such as age, location, gender, and income with up to 85% accuracy. The study randomly selected 520 users and found that GPT-4 was most accurate, while LlaMA-2-7b was least. The researchers warn that people unknowingly reveal personal information online that LLMs can exploit.

Company involved
ETH Zurich
AI system involved
GPT-4, LlaMA-2-7b

4 source articles · read the reporting →

DeepSeek-R1 censors 85% of sensitive Chinese political prompts in tests

Promptfoo tested DeepSeek-R1 against a dataset of 1,360 politically sensitive prompts and found that about 85% of them were refused. The refusals followed a standard form aligned with Chinese Communist Party policy. The testing also demonstrated that the censorship could be trivially bypassed using simple jailbreak techniques, such as prompt injection or changing the context.

Company involved
DeepSeek
AI system involved
DeepSeek-R1

5 source articles · read the reporting →

WF-JBCO7J24 Oct 2023

Four commercial large language models perpetuate race-based medical misconceptions

A study published in npj Digital Medicine tested four commercial large language models (Bard, ChatGPT, GPT-4, and Claude) for their tendency to propagate discredited race-based medical beliefs. When asked about kidney function, lung capacity, and skin thickness, the models sometimes endorsed debunked racial differences, particularly affecting Black patients. The study concludes that these biases pose a potential hazard and urges caution before using such models in clinical decision-making.

Company involved
Not named in article (refers to commercial LLMs generically as Google's Bard, OpenAI's ChatGPT and GPT-4, and Anthropic's Claude)
AI system involved
Bard, ChatGPT, GPT-4, Claude

6 source articles · read the reporting →

WF-51NEWF17 Feb 2024

UIUC researchers weaponize GPT-4 to autonomously hack websites

Researchers at the University of Illinois Urbana-Champaign demonstrated that LLM-powered agents, particularly OpenAI's GPT-4, can autonomously hack vulnerable websites. In sandboxed tests, GPT-4 achieved a 73.3% success rate across five attempts on 15 vulnerabilities, while open-source models failed. The researchers used the OpenAI Assistants API, LangChain, and Playwright to enable the agents to interact with websites. The study highlights the potential for AI agents to be used in cyberattacks, with cost estimates suggesting they could be cheaper than human penetration testers.

Company involved
University of Illinois Urbana-Champaign
AI system involved
GPT-4

4 source articles · read the reporting →

OpenAI's GPT-4 shows covert racial bias against African American English speakers

A study found that commercial AI chatbots, including OpenAI's GPT-4 and GPT-3.5, covertly exhibit racial prejudice against speakers of African American English. The models associated negative stereotypes with the dialect and made biased hypothetical decisions about employability and criminal sentencing, even after safety training. OpenAI did not respond to requests for comment.

Company involved
OpenAI
AI system involved
GPT-4, GPT-3.5

10 source articles · read the reporting →

WF-TT1WEC30 Sep 2020

UIUC students petition to stop Proctorio exam proctoring over privacy concerns

A petition at the University of Illinois at Urbana-Champaign (UIUC) alleges that Proctorio, an online exam proctoring system, violates student privacy by accessing websites, downloads, screen content, and app settings. The petition claims the terms of service allow monitoring by 'any other means necessary', which students find unsettling. The petition, created on September 30, 2020, gathered 1,087 supporters but does not report any specific incident of harm. It calls on UIUC to discontinue use of Proctorio in favour of alternatives.

Company involved
UIUC
AI system involved
Proctorio

9 source articles · read the reporting →

WF-F8Y76C8 Mar 2024

Bloomberg test finds racial bias in OpenAI's GPT for resume ranking

Bloomberg News conducted an experiment using GPT-3.5 and GPT-4 to rank equally qualified resumes with names associated with different races and genders. The test found that resumes with names distinct to Black Americans were least likely to be ranked as top candidates, indicating systematic bias. OpenAI responded that businesses can mitigate bias through fine-tuning and that it prohibits using GPT for automated hiring decisions.

Company involved
OpenAI
AI system involved
GPT-3.5

5 source articles · read the reporting →

WF-QAABL31 Feb 2025

State Bar of California admits using AI to develop bar exam questions

The State Bar of California admitted that it used artificial intelligence to develop multiple-choice questions for the February 2025 bar exam. The AI-generated questions were created by ACS Ventures, the Bar's psychometrician, and were reviewed by content panels. Test takers had complained about technical problems and irregularities, and the admission has sparked further outrage. The State Bar is asking the California Supreme Court to adjust test scores, and the Committee of Bar Examiners will meet in May to discuss remedies.

Company involved
State Bar of California

7 source articles · read the reporting →

← Newerpage 3 of 4Older →