CanLII sues Caseway AI for scraping legal database
The Canadian Legal Information Institute (CanLII) has filed a lawsuit in British Columbia Supreme Court against Caseway AI, alleging that the company's AI chatbot scraped approximately 3.5 million records from CanLII's database in bulk, violating its terms of service and copyright. CanLII claims it adds value to public court records through hyperlinks and corrections, which it says constitute protected copyrighted work. Caseway AI argues the information is public and accessible elsewhere, and that it did not use CanLII's enhancements. The lawsuit was settled in March 2026, with terms undisclosed.
- Company involved
- Canadian Legal Information Institute (CanLII)
- AI system involved
- Caseway
5 source articles · read the reporting →
Landberg v City of New York (CA NY (2d)): AI-hallucinated content in court filing, Monetary Sanction
AI-generated fake legal citations were included in a court filing, resulting in monetary sanctions against the attorney and his law firm.
1 source article · read the reporting →
Adobe uses Creative Cloud user data to train AI
Adobe's content analysis feature may scan Creative Cloud and Document Cloud files to train machine learning models, including object recognition in Lightroom and Liquid Mode in Acrobat. The company says it may analyze images, audio, video, text and other documents stored on its servers. Adobe offers an opt-out in account privacy settings, but the article notes the company did not ask first.
- Company involved
- Adobe
- AI system involved
- Content analysis
10 source articles · read the reporting →
Thomas Raynard James v. Detective Kevin Conley, et al. (S.D. Florida): AI-hallucinated content in court filing, Bar Referral
AI generated hallucinated content in a court filing, affecting the legal process and the lawyers who filed it.
1 source article · read the reporting →
Stable Diffusion amplifies racial and gender stereotypes in generated images
An analysis by Bloomberg of over 5,000 images generated by Stability AI's Stable Diffusion found that the text-to-image model amplifies racial and gender stereotypes. The model overrepresented lighter-skinned men in high-paying jobs and darker-skinned people in low-paying jobs, and underrepresented women in positions of power. Stability AI acknowledged the inherent biases in its models and stated it is working on mitigation.
- Company involved
- Stability AI
- AI system involved
- Stable Diffusion
8 source articles · read the reporting →
In re Rosslyn2016, LLC, et al. (S.D. Texas (Bankruptcy)): AI-hallucinated content in court filing, CLE on generative AI; Civil Contempt;…
The AI generated fabricated legal citations that were submitted to the bankruptcy court.
1 source article · read the reporting →
Henry v. Long Island University (E.D. New York): AI-hallucinated content in court filing, Two-year filing-disclosure sanction
The court granted a motion to dismiss and imposed sanctions on plaintiff's counsel for submitting a brief with AI-hallucinated fake cases.
1 source article · read the reporting →
Quinteros v. Harbor Distributing, LLC (CA California (1d)): AI-hallucinated content in court filing, Monetary Sanction; Bar referral
AI generated hallucinated legal citations that were filed in court, misleading the court and opposing counsel.
- Company involved
- Lipeles Law Group
1 source article · read the reporting →
JMOR Properties v. Artist Alley Townhomes et al. (CA Florida (4d)): AI-hallucinated content in court filing, Bar Referral
The AI generated fabricated legal citations that were included in a court filing, affecting the court and the opposing party.
1 source article · read the reporting →
Southland Homes & Real Estate and Investment, LLC v. Lam (SC California): AI-hallucinated content in court filing, Monetary Sanctions; Bar…
The AI system generated legal briefs containing fabricated case citations, which were filed in court, affecting the court and the opposing party.
1 source article · read the reporting →
LINAGORA closes Lucie 7B after user mockery
LINAGORA, a French open-source software company, launched a beta version of its large language model Lucie 7B. The model was intended to be a transparent and ethical alternative to big tech AI. However, after users tested it and highlighted its shortcomings, the model was mocked online. LINAGORA subsequently closed the platform to address the issues and collect more data.
- Company involved
- LINAGORA
- AI system involved
- Lucie 7B
6 source articles · read the reporting →
Stable Diffusion reproduces exact copies of training images
Researchers found that Stable Diffusion, an AI image generation model, can reproduce exact copies of images from its training dataset, including copyrighted material and personal photos. The model memorized over a thousand training examples, posing copyright and privacy risks. The researchers warn that this is an industry-wide problem affecting models like DALL-E 2 and Imagen.
- Company involved
- Stability AI
- AI system involved
- Stable Diffusion
7 source articles · read the reporting →
Authors sue Nvidia for copyright infringement in NeMo AI training
Three authors have filed a proposed class action against Nvidia, alleging the company used their copyrighted books without permission to train its NeMo AI platform. The dataset, containing approximately 196,640 volumes, was removed in October 2023 after copyright infringement reports. The authors are seeking unspecified damages on behalf of US writers whose works were used in the past three years.
- Company involved
- Nvidia
- AI system involved
- NeMo
9 source articles · read the reporting →
Anti-piracy group takes down Books3 dataset used to train Meta's LLaMA
The Danish anti-piracy group Rights Alliance sent a DMCA takedown request to The Eye, which hosted the Books3 dataset containing 196,640 copyrighted books. The dataset was used by Meta to train its LLaMA language model. Authors including Sarah Silverman have filed a class action lawsuit against Meta for using their works without permission. The dataset has been taken offline, but copies remain available.
- Company involved
- Meta
- AI system involved
- LLaMA
10 source articles · read the reporting →
Duke University recorded students' faces without proper consent for public dataset
In March 2014, Duke University researchers recorded thousands of students walking to class on campus without their knowledge or proper consent, creating the DukeMTMC dataset of over 2 million image frames. The dataset was placed on a public website and downloaded by academics, security contractors, and military researchers globally, including Chinese companies and military academies linked to surveillance of ethnic minorities. The university took down the public website in April 2019 after an Institutional Review Board investigation found the study deviated significantly from the approved protocol. The lead researcher apologized, stating he took full responsibility for his mistakes.
- Company involved
- Duke University
- AI system involved
- DukeMTMC
10 source articles · read the reporting →
Audit of LAION-400M finds sexual violence, racial slurs, and stereotypes in dataset
An audit of the LAION-400M dataset by Abeba Birhane and colleagues at University College Dublin and University of Edinburgh found that its automated curation using CLIP failed to remove sexually explicit images, racial slurs, and stereotypes. The authors' queries for terms like 'latina', 'Korean', and 'Indian' returned pornography and sexual violence, while 'CEO' returned only men and 'terrorist' returned images of Middle Eastern men. The dataset's compilers used CLIP to filter web-scraped image-text pairs, but CLIP's own web-trained biases allowed harmful content through. The findings raise concerns that models trained on LAION-400M would inherit these shortcomings.
- Company involved
- LAION-400M team
- AI system involved
- LAION-400M
7 source articles · read the reporting →
Google Lens AI overviews share misleading information about images
Google Lens's AI overviews provided false and misleading information about images, including miscaptioned videos and AI-generated footage. Full Fact found that the overviews repeated debunked claims and failed to identify inauthentic content. Google acknowledged the errors and said they were caused by problems with visual search results.
- Company involved
- Google
- AI system involved
- Google Lens AI overviews
2 source articles · read the reporting →
Luma's Dream Machine generates video with Disney character
Luma's AI video tool Dream Machine generated a trailer that included a recognizable character from Disney's Monsters, Inc. The company's CEO said a user uploaded an image containing the character, which the system then animated. The incident raises concerns about lack of transparency in training data and potential copyright infringement. Disney has not publicly commented.
- Company involved
- Luma
- AI system involved
- Dream Machine
7 source articles · read the reporting →
Developer iperov releases DeepFaceLive real-time face-swap AI on GitHub
The developer iperov has published DeepFaceLive, a neural network for real-time face swapping, on GitHub. The tool automatically replaces a user's face in live streams and video calls with a nonexistent model or a celebrity, and the installation instructions are simple. The developer claims 95% of deepfakes on YouTube were made with the related DeepFaceLab. No specific harm is reported, but the article highlights the tool's potential for misuse.
- Company involved
- iperov
- AI system involved
- DeepFaceLive
7 source articles · read the reporting →
BREIN takes down AI dataset for copyright infringement
Stichting BREIN took down a large Dutch-language dataset used to train AI models after discovering it contained illegal copies of tens of thousands of books, news articles, and subtitles. The dataset was compiled from copyrighted material without permission. The maker of the dataset signed a declaration promising not to infringe and provided information about recipients. BREIN is investigating which AI models used the dataset.
8 source articles · read the reporting →
Samsung settles Texas lawsuit over ACR data collection on smart TVs
Samsung has settled a lawsuit with the Texas Attorney General over its Automated Content Recognition (ACR) system on smart TVs. The system collected viewing data from users without informed consent. As part of the settlement, Samsung agreed to stop collecting ACR data from Texans without explicit consent and to rewrite its privacy prompts. Samsung also faces a federal class action in New York over similar allegations.
- Company involved
- Samsung
- AI system involved
- Automated Content Recognition (ACR)
7 source articles · read the reporting →
Stable Diffusion Accused of Stealing Artists' Styles Without Consent
Artists including Greg Rutkowski and Karla Ortiz allege that Stability AI's image generator Stable Diffusion was trained on their work without permission, enabling users to create images mimicking their distinctive styles. The artists express concern that this threatens their livelihoods and identities, as their names become prompts for generating similar art. A tool is being developed to help protect artists from such unauthorised use, but the situation remains unresolved.
- Company involved
- Stability AI
- AI system involved
- Stable Diffusion
10 source articles · read the reporting →
Duke University MTMC Dataset Used in Authoritarian Surveillance Research
Duke University created and openly distributed the Duke MTMC dataset, containing surveillance footage of approximately 2,000 students and visitors on campus. The dataset was used by numerous organisations, including Chinese military-linked companies like SenseTime and Hikvision, for developing person re-identification and facial recognition technologies. Following an investigation by exposing.ai and the Financial Times, Duke University terminated the dataset in May 2019. The incident highlights the privacy risks of academic datasets being repurposed for mass surveillance without consent.
- Company involved
- Duke University
- AI system involved
- Duke MTMC
1 source article · read the reporting →
Adobe Firefly trained on thousands of Midjourney images, Bloomberg reports
Bloomberg has reported that Adobe's Firefly image generator was trained using thousands of images from competitor Midjourney. Adobe says these made up about 5% of the training data and were part of the Adobe Stock library. The company has marketed Firefly as ethically trained and offered enterprise customers indemnity against copyright claims. Adobe responded that all Adobe Stock images undergo moderation, but the report has raised questions about Firefly's copyright safety.
- Company involved
- Adobe
- AI system involved
- Firefly
7 source articles · read the reporting →