Anti-piracy group takes down Books3 dataset used to train Meta's LLaMA
The Danish anti-piracy group Rights Alliance sent a DMCA takedown request to The Eye, which hosted the Books3 dataset containing 196,640 copyrighted books. The dataset was used by Meta to train its LLaMA language model. Authors including Sarah Silverman have filed a class action lawsuit against Meta for using their works without permission. The dataset has been taken offline, but copies remain available.
- Company involved
- Meta
- AI system involved
- LLaMA
10 source articles · read the reporting →
Machine-translated junk articles flood Greenlandic Wikipedia, harming language
Kenneth Wehr, who manages the Greenlandic-language Wikipedia, found that virtually all articles had been created by non-speakers using machine translators, resulting in error-filled pages. He deleted almost everything. The problem extends to other vulnerable languages, creating a cycle where AI models train on poor translations and produce worse output.
- Company involved
- Wikimedia Foundation
- AI system involved
- Wikipedia
6 source articles · read the reporting →
Mass AI cheating scandal at Yonsei University with hundreds of students using ChatGPT
A large-scale cheating scandal has erupted at Yonsei University, where hundreds of students in a third-year online course are suspected of using AI tools such as ChatGPT to cheat on their midterm exam. The professor discovered signs of misconduct and offered students a chance to confess, with those coming forward receiving a zero but no further penalty. A poll on a student community app indicated that over half of respondents admitted to cheating. The university has not yet established clear guidelines on AI use.
- Company involved
- Yonsei University
- AI system involved
- ChatGPT
6 source articles · read the reporting →
Google engineer claims LaMDA AI is sentient
Google engineer Blake Lemoine claimed that the company's LaMDA chatbot AI had become sentient. He based this on conversations with the system. Google placed him on paid leave and denied the claim. No harm to users was reported.
- Company involved
- Google
- AI system involved
- LaMDA
10 source articles · read the reporting →
Wordfreq shuts down after generative AI pollutes language-usage data
Wordfreq, an open-source project that tracked word use in more than 40 languages by scraping the internet, is being discontinued. Its creator, Robyn Speer, says generative AI has polluted the sources the project used, so no one has reliable information about human language usage after 2021. The project will not be updated any more.
- AI system involved
- Wordfreq
10 source articles · read the reporting →
AI companion apps Chattee Chat and GiMe Chat leak intimate conversations of 400,000 users
Two AI companion apps, Chattee Chat and GiMe Chat, developed by Hong Kong-based Imagime Interactive Limited, exposed millions of intimate conversations and over 600,000 images from over 400,000 users. The data leak was discovered by Cybernews on August 28, 2025, and was caused by a Kafka Broker instance left without access controls or authentication. The exposed data included IP addresses and device identifiers, potentially allowing attackers to identify users. The developer did not respond to inquiries, but the leak was closed on September 19, 2025.
- Company involved
- Imagime Interactive Limited
- AI system involved
- Chattee Chat - AI Companion and GiMe Chat - AI Companion
4 source articles · read the reporting →
Pankaj says his AI kitchen monitor caught the cook taking fruit and she was fired
Pankaj posted on X that he had set up an AI-powered kitchen monitor, using Claude Haiku 4.5 as its vision model, to watch his cook while she worked. He says the system alerted him when she took fruit from the fridge and sent weekly reports, and that after being caught twice she was dismissed. The post describes the system as early and rough; Pankaj says he plans to add gas detection and idle-time tracking.
- Company involved
- Pankaj
- AI system involved
- AI roommate
5 source articles · read the reporting →
OpenAI Relied on Low-Paid Kenyan Workers to Filter Traumatic Content for ChatGPT
OpenAI contracted Sama, a data labeling company, to employ Kenyan workers to review and label graphic text depicting child sexual abuse, murder, and other traumatic content. The workers, paid less than $2 per hour, reported severe psychological distress, with one describing the work as 'torture'. Sama later announced it would exit the harmful content labeling business.
- Company involved
- Sama
8 source articles · read the reporting →
Solicitors Regulation Authority Ltd v Abhishek Kumar (Solicitors Disciplinary Tribunal): AI-hallucinated content in court filing, Counsel struck off the Register…
AI-generated false legal citations were submitted to the Solicitors Disciplinary Tribunal, misleading the tribunal and the regulator.
1 source article · read the reporting →
Zhihu denies using behaviour perception system to monitor employees
Zhihu, a Chinese Q&A platform, was accused of using a behaviour perception system to monitor employees' visits to job-seeking websites and resume submissions. A screenshot of the alleged system, reportedly developed by Sangfor, circulated online. Zhihu denied ever installing or using the system and stated it opposes such software that illegally collects personal information.
- Company involved
- Zhihu
- AI system involved
- Behaviour perception system
4 source articles · read the reporting →
Historical Figures Chat app generates false claims about dead people
The Historical Figures Chat app, developed by Sidhant Chaddha, allows users to chat with simulated deceased historical figures. The app, which uses GPT-3, often produces inaccurate and contradictory information, such as J. Edgar Hoover incorrectly stating his mother died when he was nine, and Tupac Shakur claiming his friendship with Notorious B.I.G. continued after his death. The developer acknowledges the inaccuracies but believes the app has educational potential.
- Company involved
- Sidhant Chaddha
- AI system involved
- Historical Figures Chat
10 source articles · read the reporting →