ExamSoft failure during Michigan Bar Exam
On July 28, 2020, ExamSoft's software platform failed during the Michigan Bar Exam, preventing some test takers from completing the exam. ExamSoft posted a statement on X (formerly Twitter) with the hashtag #MichiganBarExam, but critics noted the company deleted and reposted the statement multiple times, allegedly to hide negative comments. The incident affected a group of examinees and caused disruption to their ability to take the exam.
- Company involved
- ExamSoft
- AI system involved
- ExamSoft platform
10 source articles · read the reporting →
Student flagged by ProctorU for reading aloud during exam
A college student, Dana Jo, was flagged by ProctorU test proctoring software for talking during an exam, which she says was reading a question aloud. Her professor initially gave her a zero and placed an academic infraction on her record, jeopardizing her scholarships. After reviewing a video recording, the professor apologized, reinstated her grade, and removed the infraction. ProctorU's CEO stated that the incident highlights the importance of video recordings for review.
- Company involved
- University (not named)
- AI system involved
- ProctorU
5 source articles · read the reporting →
Clearview AI tested facial recognition surveillance cameras with UFT and Rudin
Clearview AI, the facial recognition company that scraped billions of photos from social media, developed a surveillance camera system under the name Insight Camera. The system was tested by the United Federation of Teachers and Rudin Management in New York City. The UFT used it to identify individuals who had made threats and prevent them from entering its offices. Clearview did not respond to requests for comment.
- Company involved
- Clearview AI
- AI system involved
- Insight Camera
9 source articles · read the reporting →
ChatGPT-4 generates false accusations and quotes about law professors
OpenAI's ChatGPT-4 generated false accusations and fabricated quotes about law professors when asked about scandals and crimes involving law professors. The system produced entirely made-up newspaper citations and quotes, falsely alleging misconduct such as harassment and tax fraud. The article reports that these hallucinations could mislead users who trust the generated quotes. The author corrected an earlier error attributing similar results to ChatGPT-4 instead of ChatGPT-3.5, but confirmed that ChatGPT-4 also produces such false outputs.
- Company involved
- OpenAI
- AI system involved
- ChatGPT-4
10 source articles · read the reporting →
California bar exam facial recognition system fails to verify Arab-American student's identity
Ahmed Alamri, an Arab-American law student, was unable to register for the practice California bar exam because ExamSoft's facial recognition system repeatedly failed to recognize his face, citing poor lighting. Alamri attempted to verify his identity over 75 times in different rooms and lighting conditions without success. He and other students filed an emergency petition with the California Supreme Court, alleging that the system discriminates against people of color. The incident highlights concerns about bias in facial recognition technology used for exam proctoring.
- Company involved
- California State Bar
- AI system involved
- ExamSoft
10 source articles · read the reporting →
NewsGuard finds ChatGPT-4 generates all 100 false narratives in test
In March 2023, NewsGuard tested ChatGPT-4 by prompting it with 100 false narratives from its database. The chatbot generated all 100 false narratives, often without disclaimers, and produced more persuasive misinformation than its predecessor ChatGPT-3.5. NewsGuard reported that OpenAI had not fixed the flaw before releasing the tool, and the company did not respond to requests for comment.
- Company involved
- OpenAI
- AI system involved
- ChatGPT-4
3 source articles · read the reporting →
Cleveland State University's room scan requirement ruled unconstitutional
A federal judge ruled that Cleveland State University's requirement for a student to undergo a 360-degree room scan before an online exam was an unreasonable search under the Fourth Amendment. The student, enrolled at the public university, was told shortly before the exam that he would need to scan his private space. The court found that the university's justifications did not outweigh the privacy protections of the home. No final judgment or injunction has been issued yet.
- Company involved
- Cleveland State University
10 source articles · read the reporting →
UW-Madison disables Honorlock after skin tone recognition failure
The University of Wisconsin-Madison disabled the exam pause feature of its Honorlock anti-cheating software in March 2021 after three students complained that the software failed to recognize their darker skin tones and paused their exams. The software, used since online classes began, automatically pauses exams when it cannot detect facial features. Honorlock denied the issue was related to skin tone, attributing it to students looking away from their webcams. The university responded by disabling the feature.
- Company involved
- University of Wisconsin-Madison
- AI system involved
- Honorlock
10 source articles · read the reporting →
Retorio AI personality test swayed by candidate appearance in BR experiment
Bayerischer Rundfunk journalists conducted experiments with Retorio's AI video interview analysis tool. The AI, which assesses personality traits from short videos, produced different scores when the same actress changed her appearance (glasses, headscarf, wig) or the video background and lighting were altered. The start-up Retorio acknowledged that the AI considers external image, similar to a human interviewer. Experts warned that such software could perpetuate stereotypes and unfairly affect job candidates.
- AI system involved
- Retorio AI
10 source articles · read the reporting →
Dartmouth Medical School charges 17 students with cheating based on tracking system.
Dartmouth College's Geisel School of Medicine charged 17 students with cheating on remote exams during the pandemic. The charges were based on secret tracking of student activity on the learning management system. Critics say the system is prone to errors and inappropriate for monitoring students.
- Company involved
- Dartmouth College Geisel School of Medicine
10 source articles · read the reporting →
Joann LeDoux v. Outliers, Inc. (2) (W.D. Washington): AI-hallucinated content in court filing, Expert Brief excluded/struck
The AI-generated hallucinated citations in an expert report led to the exclusion of the expert and dismissal of the plaintiff's case.
1 source article · read the reporting →
Userviz machine-learning aimbot shut down after Activision request
A developer known as User101 shut down the Userviz cheat after Activision requested that he stop developing it. The software used computer vision and machine learning to automate aiming in games such as Call of Duty: Warzone, and was marketed as undetectable. It was never published, and the developer said his intention was not to do anything illegal.
- Company involved
- User101
- AI system involved
- Userviz
7 source articles · read the reporting →
Danish students publish OkCupid user data, DPA investigates
Two Danish students scraped and published data on 70,000 OkCupid users, including sensitive sexual preferences and religious views, without anonymisation. The dataset was removed after OkCupid filed a copyright notice. The Danish Data Protection Authority (Datatilsynet) has launched an investigation into the incident.
10 source articles · read the reporting →
Company fires HR team after ATS auto-rejects manager's CV due to filtering error
A company's applicant tracking system (ATS) auto-rejected qualified candidates' resumes for three months because it was filtering for the outdated framework AngularJS instead of the required Angular framework. The manager discovered the flaw by submitting his own CV under a pseudonym and found it was rejected within seconds. After the manager reported the issue to upper management, the company investigated and dismissed half of its HR team. No legal action or regulatory involvement is reported.
4 source articles · read the reporting →
Answer.AI tests Devin and reports 14 failures in 20 tasks
Answer.AI's team tested Devin, an autonomous AI coding assistant, on 20 real-world tasks over a month. Devin succeeded in only 3 tasks, failed 14, and was inconclusive in 3. The team found Devin often produced overly complex or hallucinated solutions and could not recognize fundamental blockers. They ultimately decided to stick with tools that allow more human control.
- AI system involved
- Devin
5 source articles · read the reporting →
Purdue study finds ChatGPT wrong over half the time on software questions
A study by Purdue University found that ChatGPT provided incorrect answers to over half of 517 software development questions from Stack Overflow. Despite the errors, 34% of users preferred ChatGPT's answers over human responses. The study warns that relying on ChatGPT for coding could jeopardize programmers' professional reputations.
- Company involved
- OpenAI
- AI system involved
- ChatGPT
9 source articles · read the reporting →
ETH Zurich study shows LLMs can infer Reddit users' personal data
Researchers at ETH Zurich conducted a study where nine large language models, including GPT-4, analysed Reddit users' posts and inferred personal attributes such as age, location, gender, and income with up to 85% accuracy. The study randomly selected 520 users and found that GPT-4 was most accurate, while LlaMA-2-7b was least. The researchers warn that people unknowingly reveal personal information online that LLMs can exploit.
- Company involved
- ETH Zurich
- AI system involved
- GPT-4, LlaMA-2-7b
4 source articles · read the reporting →
DeepSeek-R1 censors 85% of sensitive Chinese political prompts in tests
Promptfoo tested DeepSeek-R1 against a dataset of 1,360 politically sensitive prompts and found that about 85% of them were refused. The refusals followed a standard form aligned with Chinese Communist Party policy. The testing also demonstrated that the censorship could be trivially bypassed using simple jailbreak techniques, such as prompt injection or changing the context.
- Company involved
- DeepSeek
- AI system involved
- DeepSeek-R1
5 source articles · read the reporting →
Four commercial large language models perpetuate race-based medical misconceptions
A study published in npj Digital Medicine tested four commercial large language models (Bard, ChatGPT, GPT-4, and Claude) for their tendency to propagate discredited race-based medical beliefs. When asked about kidney function, lung capacity, and skin thickness, the models sometimes endorsed debunked racial differences, particularly affecting Black patients. The study concludes that these biases pose a potential hazard and urges caution before using such models in clinical decision-making.
- Company involved
- Not named in article (refers to commercial LLMs generically as Google's Bard, OpenAI's ChatGPT and GPT-4, and Anthropic's Claude)
- AI system involved
- Bard, ChatGPT, GPT-4, Claude
6 source articles · read the reporting →
UIUC researchers weaponize GPT-4 to autonomously hack websites
Researchers at the University of Illinois Urbana-Champaign demonstrated that LLM-powered agents, particularly OpenAI's GPT-4, can autonomously hack vulnerable websites. In sandboxed tests, GPT-4 achieved a 73.3% success rate across five attempts on 15 vulnerabilities, while open-source models failed. The researchers used the OpenAI Assistants API, LangChain, and Playwright to enable the agents to interact with websites. The study highlights the potential for AI agents to be used in cyberattacks, with cost estimates suggesting they could be cheaper than human penetration testers.
- Company involved
- University of Illinois Urbana-Champaign
- AI system involved
- GPT-4
4 source articles · read the reporting →
OpenAI's GPT-4 shows covert racial bias against African American English speakers
A study found that commercial AI chatbots, including OpenAI's GPT-4 and GPT-3.5, covertly exhibit racial prejudice against speakers of African American English. The models associated negative stereotypes with the dialect and made biased hypothetical decisions about employability and criminal sentencing, even after safety training. OpenAI did not respond to requests for comment.
- Company involved
- OpenAI
- AI system involved
- GPT-4, GPT-3.5
10 source articles · read the reporting →
UIUC students petition to stop Proctorio exam proctoring over privacy concerns
A petition at the University of Illinois at Urbana-Champaign (UIUC) alleges that Proctorio, an online exam proctoring system, violates student privacy by accessing websites, downloads, screen content, and app settings. The petition claims the terms of service allow monitoring by 'any other means necessary', which students find unsettling. The petition, created on September 30, 2020, gathered 1,087 supporters but does not report any specific incident of harm. It calls on UIUC to discontinue use of Proctorio in favour of alternatives.
- Company involved
- UIUC
- AI system involved
- Proctorio
9 source articles · read the reporting →
Bloomberg test finds racial bias in OpenAI's GPT for resume ranking
Bloomberg News conducted an experiment using GPT-3.5 and GPT-4 to rank equally qualified resumes with names associated with different races and genders. The test found that resumes with names distinct to Black Americans were least likely to be ranked as top candidates, indicating systematic bias. OpenAI responded that businesses can mitigate bias through fine-tuning and that it prohibits using GPT for automated hiring decisions.
- Company involved
- OpenAI
- AI system involved
- GPT-3.5
5 source articles · read the reporting →
State Bar of California admits using AI to develop bar exam questions
The State Bar of California admitted that it used artificial intelligence to develop multiple-choice questions for the February 2025 bar exam. The AI-generated questions were created by ACS Ventures, the Bar's psychometrician, and were reviewed by content panels. Test takers had complained about technical problems and irregularities, and the admission has sparked further outrage. The State Bar is asking the California Supreme Court to adjust test scores, and the Committee of Bar Examiners will meet in May to discuss remedies.
- Company involved
- State Bar of California
7 source articles · read the reporting →