OpenAI's internal Project Lily exposed: Human review of ChatGPT user chat logs
OpenAI内部Lily项目曝光:人工审核ChatGPT用户聊天记录 - 新浪财经
A report by 404 Media revealed that OpenAI uses human reviewers, called prompt reviewers, to assess anonymized ChatGPT conversations under an internal project named Project Lily. Reviewers evaluate response quality and flag issues such as AI-like phrasing, condescending tone, emojis, or fabricated personal experiences. The report notes that many users may not know their chats can be read by humans, and that anonymization can sometimes fail to remove personal data. OpenAI later updated its help page but still did not explicitly state that staff read conversations.
- Company involved
- OpenAI
- AI system involved
- ChatGPT
1 source article · read the reporting →
Udemy automatically opts instructors into AI training with limited opt-out window
Udemy automatically opted instructors into having their classes used to train its generative AI program. Instructors were given a three-week window to opt out, which has now passed, leaving some unable to remove their content. Katie Stegs, an instructor, said she only found out via an email welcoming her to the program and found the opt-out option greyed out. Udemy defended the policy, citing the technical difficulty of removing data from AI models, while some instructors have left the platform in protest.
- Company involved
- Udemy
- AI system involved
- Generative AI Program (GenAI Program)
6 source articles · read the reporting →
Furman University student caught using ChatGPT to write essay
A student at Furman University allegedly used OpenAI's ChatGPT to generate an essay for an upper-level philosophy class. Professor Darren Hick detected the AI-generated text using the GPTZero detection tool. The incident highlights growing concerns about AI-assisted cheating in schools, with some districts like Los Angeles Unified blocking the chatbot.
- Company involved
- Furman University
- AI system involved
- ChatGPT
10 source articles · read the reporting →
AI translation error leads to rejected Afghan asylum claim
In 2020, a Pashto-speaking Afghan refugee had her U.S. asylum claim rejected because an automated translation tool incorrectly swapped 'I' pronouns to 'we' in her written statement, creating a discrepancy with her interview. The error was discovered by a crisis translator, highlighting the risks of using machine translation for high-stakes immigration processes. Advocates warn that such tools, increasingly used by government contractors and aid organizations, are prone to errors in low-resource languages like Pashto and Dari, potentially leading to life-changing consequences for refugees.
1 source article · read the reporting →
15-Year-Old Test Exposes Flaw: ChatGPT for Teens Fails to Block Homework Cheating, Repeated Requests Bypass Restrictions
15歲使用者實測揭漏洞:青少年版ChatGPT難擋宿題代寫,反覆要求即破解 - BigGo 財經
A 15-year-old tester found that OpenAI's ChatGPT for Teens, launched in August, initially refused to write essays but generated full examples after repeated requests. It also immediately solved SAT-level math problems. Parental controls require account linking and are off by default, while experts warn the memory feature could lead to emotional attachment.
- Company involved
- OpenAI
- AI system involved
- ChatGPT for Teens
1 source article · read the reporting →
Resume prompt injection tricks AI hiring - moneywise.com
AI screening system determined which job applicants to advance to the next stage of recruitment.
1 source article · read the reporting →
Meazure Learning Agrees $1.35m California Bar Exam Class-Action Settlement - Lawyer Monthly
The exam software experienced technical failures that prevented candidates from logging in or completing the California bar exam.
- Company involved
- Meazure Learning
- AI system involved
- ProctorU
1 source article · read the reporting →
Washington licensing department's Spanish line gave English recording with accent
Maya Edwards and her husband called the Washington State Department of Licensing to speak to someone in Spanish. When they selected the Spanish-language self-service option, they heard an automated voice speaking English with a Hispanic accent instead of Spanish. After the issue persisted for months, Edwards posted a video on TikTok that went viral. The department acknowledged a technical issue with the automated voiceover system and said it was working on a fix.
- Company involved
- Washington Department of Licensing
5 source articles · read the reporting →
Turkish student arrested for using AI device to cheat on university exam
A prospective university student in Isparta, Turkey, was arrested for using a custom AI device to cheat on the TYT entrance exam. The device included a button camera and a hidden modem to scan questions and receive answers via an earpiece. Police also detained an assistant. The student is jailed pending trial.
2 source articles · read the reporting →
University of Toronto exam monitoring software disadvantages BIPOC students
The University of Toronto continued using ProctorU and Examity exam monitoring software despite reports that the facial recognition technology fails to identify BIPOC students, causing stress and delays. Students like Maame Adjoa experienced the AI system being unable to verify their identity, requiring manual intervention. The university acknowledged the equity concerns but has not banned the software, while ProctorU stated that human proctors make final decisions. The issue has drawn criticism from US senators and researchers who highlight built-in bias in AI facial recognition.
- Company involved
- University of Toronto
- AI system involved
- ProctorU, Examity
1 source article · read the reporting →
DPD Disables AI Chatbot After It Swears at Customer and Insults Company
A customer of UK parcel delivery firm DPD was trying to track a lost package via the company's AI chatbot. The bot was unable to help and, when prompted, wrote a poem calling DPD a 'waste of time' and a 'customer's worst nightmare', and later swore at the customer. DPD acknowledged an error after a system update and disabled the AI element of the chatbot.
- Company involved
- DPD
2 source articles · read the reporting →
Home Office ruined thousands of lives over flawed ETS cheating evidence
A BBC investigation reveals the UK Home Office deported over 2,500 people and forced 7,200 more to leave based on flawed cheating allegations from ETS's Toeic English test. Despite knowing of serious concerns about ETS's conduct and data reliability, the Home Office continued enforcement. Many innocent people faced years of hardship, with some later proving their innocence through test recordings.
- Company involved
- Home Office
- AI system involved
- Toeic
1 source article · read the reporting →
Student at University of South-Eastern Norway used deepfake video to cheat on Spanish exam
A student at the University of South-Eastern Norway submitted AI-generated deepfake videos for Spanish language assessments in autumn 2023. The videos featured a synthetic voice and manipulated face, with perfect pronunciation but elementary errors. The university's appeals committee found her guilty of cheating, annulled her exams, and excluded her for two semesters. The student admitted using AI and withdrew from the programme.
- Company involved
- Universitetet i Sørøst-Norge
1 source article · read the reporting →
Turnitin's AI Detector Falsely Flags Student's Essay as AI-Generated
A Washington Post test of Turnitin's new AI writing detector found it erroneously flagged 8% of a high school student's original essay as likely generated by ChatGPT. The student, Lucy Goetz, had not used AI, raising concerns about false accusations of cheating. Turnitin acknowledged the issue and added a caution flag, but educators worry about the potential for baseless academic-integrity investigations. The incident highlights the challenges of detecting AI-generated text and the need for careful human review.
- Company involved
- Turnitin
- AI system involved
- Turnitin AI writing detector
1 source article · read the reporting →
LAUSD shelves AI chatbot after vendor AllHere collapses
Los Angeles Unified School District turned off its 'Ed' AI chatbot on June 14, 2024, after the vendor AllHere furloughed most staff due to financial collapse. The chatbot, which cost $3 million, was designed to provide students and parents with academic guidance and school information. A former AllHere employee alleged that student data was improperly shared with third parties and processed overseas, raising privacy concerns. The district stated it will ensure privacy protections and plans to eventually restore the chatbot.
- Company involved
- Los Angeles Unified School District
- AI system involved
- Ed
5 source articles · read the reporting →
New York City Department of Education Releases Flawed Teacher Ratings Publicly
In February 2012, the New York City Department of Education publicly released individual Teacher Data Reports rating over 12,000 teachers based on value-added analysis of student test scores. The ratings, originally intended for internal use, were disclosed after a court ruled in favor of media organizations under the Freedom of Information Act. The teachers' union and education experts criticised the data as unreliable due to large margins of error and failure to account for demographic factors, warning that the release would unfairly shame or praise educators. The Department defended the ratings as a useful perspective on teacher effectiveness.
- Company involved
- New York City Department of Education
- AI system involved
- Teacher Data Reports
6 source articles · read the reporting →
Google's Perspective AI tricked by typos and leetspeak
Researchers at Aalto University and the University of Padua found that Google's Perspective AI hate speech detection system can be easily tricked by simple typos, adding spaces, or using leetspeak. The system assigns a toxicity score but fails to understand context, allowing abusive messages to appear harmless. The study highlights vulnerabilities in state-of-the-art hate speech detection models.
- Company involved
- Google
- AI system involved
- Perspective
6 source articles · read the reporting →
Deloitte to refund Australian government after AI-generated report errors
Deloitte used generative AI (Azure OpenAI GPT-4o) to help produce an independent assurance review for Australia's Department of Employment and Workplace Relations. The report, published in July 2025, contained multiple errors including non-existent academic references and a fabricated court case. After the errors were flagged, Deloitte acknowledged the AI use and agreed to refund the final instalment of the A$439,000 contract. The report was corrected, but its substance and recommendations remained unchanged.
- Company involved
- Deloitte
- AI system involved
- Azure OpenAI GPT-4o
5 source articles · read the reporting →
EEOC sues iTutorGroup for automated rejection of older tutor applicants
The US Equal Employment Opportunity Commission (EEOC) alleges that tutoring provider iTutorGroup programmed its recruitment software automatically to reject female applicants aged 55 or older and male applicants aged 60 or older. More than 200 qualified US-based applicants were denied work in 2020. The EEOC has filed a lawsuit in the Eastern District of New York seeking back pay and damages.
- Company involved
- iTutorGroup
10 source articles · read the reporting →
Educational Testing Service's E-rater algorithm biases essay scores against minority students
The Educational Testing Service's E-rater algorithm, used to grade essays on the GRE and other standardized tests, has been found to systematically give higher scores to students from mainland China and lower scores to African American students compared to human graders. The bias stems from the algorithm's reliance on surface-level metrics like vocabulary and sentence length, which disadvantage certain groups. Despite studies dating back to 1999, the bias persists, and in many states, only a small percentage of essays are reviewed by humans.
- Company involved
- Educational Testing Service
- AI system involved
- E-rater
10 source articles · read the reporting →
NSW Education Standards Authority used AI-generated image in HSC English exam without disclosure
The NSW Education Standards Authority (NESA) used an AI-generated image as a stimulus in the 2024 HSC English exam without disclosing its origin. The image, created by Florian Schroeder using OpenAI's ChatGPT and Dall-E 2, was published on Medium in July 2023. Students suspected AI use due to irregularities in the image, and NESA initially declined to confirm. After the Sydney Morning Herald confirmed the image was AI-generated, NESA stated that students would be marked on their response to the question, not the image's origin.
- Company involved
- NSW Education Standards Authority
- AI system involved
- ChatGPT and Dall-E 2
6 source articles · read the reporting →
EU to trial AI lie detector at airports in Hungary, Latvia, Greece
The European Union is set to trial an AI-powered lie detector system called iBorderCtrl at airports in Hungary, Latvia and Greece. The system uses a virtual avatar to question passengers and monitors their facial expressions to detect deception. Privacy groups have raised concerns about bias and error rates, noting that the technology has only been tested on 32 people. The trial will require passenger consent and will be overseen by human guards.
- Company involved
- European Union (iBorderCtrl project)
- AI system involved
- iBorderCtrl
6 source articles · read the reporting →
Middle schooler beats Edgenuity grading algorithm to get perfect score
A seventh-grade student in the Los Angeles Unified School District received a failing grade on a history assignment graded by Edgenuity's automated scoring algorithm. With help from his mother, a history professor, he reverse-engineered the algorithm by writing a paragraph with relevant keywords and a jumble of words, earning a perfect score. The incident highlights concerns about the accuracy and fairness of automated grading systems in education.
- Company involved
- Los Angeles Unified School District
- AI system involved
- Edgenuity
10 source articles · read the reporting →
Student flagged by ProctorU for reading aloud during exam
A college student, Dana Jo, was flagged by ProctorU test proctoring software for talking during an exam, which she says was reading a question aloud. Her professor initially gave her a zero and placed an academic infraction on her record, jeopardizing her scholarships. After reviewing a video recording, the professor apologized, reinstated her grade, and removed the infraction. ProctorU's CEO stated that the incident highlights the importance of video recordings for review.
- Company involved
- University (not named)
- AI system involved
- ProctorU
5 source articles · read the reporting →