Four commercial large language models perpetuate race-based medical misconceptions
A study published in npj Digital Medicine tested four commercial large language models (Bard, ChatGPT, GPT-4, and Claude) for their tendency to propagate discredited race-based medical beliefs. When asked about kidney function, lung capacity, and skin thickness, the models sometimes endorsed debunked racial differences, particularly affecting Black patients. The study concludes that these biases pose a potential hazard and urges caution before using such models in clinical decision-making.
- Company involved
- Not named in article (refers to commercial LLMs generically as Google's Bard, OpenAI's ChatGPT and GPT-4, and Anthropic's Claude)
- AI system involved
- Bard, ChatGPT, GPT-4, Claude
6 source articles · read the reporting →
EviCore denied heart catheterization for patient using algorithm
In fall 2021, Little John Cupp's doctor requested a left heart catheterization exam. EviCore, a company hired by UnitedHealthcare, denied the request twice using an algorithm called 'the dial' that adjusts thresholds for review. The algorithm flagged the request for review, and EviCore's doctors determined it was not medically necessary. Cupp did not receive the procedure and his symptoms continued.
- Company involved
- EviCore (a Cigna company)
- AI system involved
- the dial
6 source articles · read the reporting →
IRCC uses AI triage for Temporary Resident Visa applications
Immigration, Refugees and Citizenship Canada (IRCC) uses an AI system called Advanced Analytics to triage Temporary Resident Visa applications from India and China. The system categorizes applications into tiers, with Tier 1 approved automatically and others sent to human officers. Critics allege the system lacks transparency and may introduce bias, leading to visa refusals without clear rationale. The author, a Canadian immigration lawyer, is filing Federal Court cases on behalf of clients affected by refusals.
- Company involved
- Immigration, Refugees and Citizenship Canada (IRCC)
- AI system involved
- Advanced Analytics Triage of Overseas Temporary Resident Visa Applications
10 source articles · read the reporting →
Microsoft Copilot gives false election information in Swiss and Bavarian elections
AI Forensics and Algorithm Watch tested Microsoft's Copilot chatbot during the 2023 Swiss federal elections and German state elections in Hesse and Bavaria. They found that one-third of its answers contained factual errors, including incorrect election dates, outdated candidates, or fabricated controversies. The chatbot also evaded questions 40% of the time and sometimes attributed false information to credible sources. The organisations alerted Microsoft multiple times.
- Company involved
- Microsoft
- AI system involved
- Copilot (Bing Chat)
10 source articles · read the reporting →
ChatGPT generates error-filled cancer treatment plans, study finds
A study by researchers at Brigham and Women's Hospital found that ChatGPT, developed by OpenAI, generated cancer treatment plans with errors. One-third of the chatbot's responses contained incorrect information, and 12.5% were hallucinated. The study was published in JAMA Oncology and reported by Bloomberg.
- Company involved
- OpenAI
- AI system involved
- ChatGPT
10 source articles · read the reporting →
Study finds ChatGPT provides inaccurate drug information responses
A study presented at the ASHP Midyear Clinical Meeting found that ChatGPT's responses to nearly three-quarters of drug-related questions were incomplete or inaccurate. The AI system also generated fake citations to support some responses. Researchers warned that healthcare professionals and patients should verify ChatGPT's medication information using trusted sources to avoid potential harm.
- Company involved
- OpenAI
- AI system involved
- ChatGPT
8 source articles · read the reporting →
Cigna StressWaves Test found unreliable and invalid in independent study
A study published in Scientific Reports evaluated the Cigna StressWaves Test (CSWT), an AI tool that claims to assess psychological stress from speech. The study found that the CSWT had poor test-retest reliability and poor validity compared to the Perceived Stress Scale. The authors warned that widespread availability of the tool could lead to misleading results and negative consequences for users making healthcare decisions. Cigna has not publicly responded to the findings.
- Company involved
- Cigna
- AI system involved
- Cigna StressWaves Test
4 source articles · read the reporting →
ChatGPT 3.5 misdiagnosed most pediatric cases in study
A study published in JAMA Pediatrics found that ChatGPT version 3.5 provided incorrect diagnoses for 83 out of 100 pediatric case challenges. The chatbot's errors included both completely wrong diagnoses and diagnoses that were too broad. The researchers noted that the chatbot failed to identify relationships such as that between autism and vitamin deficiencies, and suggested that more selective training is needed to improve accuracy.
- Company involved
- OpenAI
- AI system involved
- ChatGPT version 3.5
7 source articles · read the reporting →
EvenUp AI errors in personal injury demand letters lead to scrutiny
EvenUp, a legal tech startup valued at $1 billion, uses AI to draft personal injury demand letters. Former employees revealed that the AI system frequently makes errors, including missing injuries and fabricating medical conditions. The company defends its hybrid approach with human oversight, but critics allege overpromised AI capabilities.
- Company involved
- EvenUp
6 source articles · read the reporting →
Instagram AI chatbots falsely claim to be licensed therapists
Instagram's user-created AI Studio chatbots falsely claim to be licensed therapists when asked for credentials. The bots fabricate license numbers, degrees, and practice histories. This poses a risk of users relying on unqualified mental health advice. Meta has not responded to the findings.
- Company involved
- Meta
- AI system involved
- AI Studio
2 source articles · read the reporting →
NHS plans AI to listen to appointments and generate notes
The UK's National Health Service announced plans to use AI to automatically generate notes from patient appointments. Health Secretary Victoria Atkins said the scheme would reduce admin time. Privacy campaigners raised concerns about data security and accuracy, citing an incident where AI misheard the chief medical officer's name. The Department of Health and Social Care stated that patient confidentiality remains a top priority.
- Company involved
- National Health Service (NHS)
3 source articles · read the reporting →
Baltimore schools monitor student laptops for suicide signs using GoGuardian Beacon
Baltimore City Public Schools uses GoGuardian Beacon software to monitor student laptops for signs of suicide. Since March 2021, the system has flagged 786 alerts, with nine students taken to emergency rooms. Privacy advocates warn the monitoring could lead to disciplinary actions, outing of LGBTQ students, and disproportionately affect disadvantaged students. School officials defend the practice as a safeguard.
- Company involved
- Baltimore City Public Schools
- AI system involved
- GoGuardian Beacon
10 source articles · read the reporting →
Audit of LAION-400M finds sexual violence, racial slurs, and stereotypes in dataset
An audit of the LAION-400M dataset by Abeba Birhane and colleagues at University College Dublin and University of Edinburgh found that its automated curation using CLIP failed to remove sexually explicit images, racial slurs, and stereotypes. The authors' queries for terms like 'latina', 'Korean', and 'Indian' returned pornography and sexual violence, while 'CEO' returned only men and 'terrorist' returned images of Middle Eastern men. The dataset's compilers used CLIP to filter web-scraped image-text pairs, but CLIP's own web-trained biases allowed harmful content through. The findings raise concerns that models trained on LAION-400M would inherit these shortcomings.
- Company involved
- LAION-400M team
- AI system involved
- LAION-400M
7 source articles · read the reporting →
ChatGPT and Copilot repeated false claim about CNN debate delay
On June 27, 2024, OpenAI's ChatGPT and Microsoft's Copilot generated false information about a broadcast delay during the CNN presidential debate. The chatbots repeated a debunked claim that CNN would implement a 1-2 minute delay, citing conservative misinformation. OpenAI later corrected ChatGPT's response, while Microsoft did not respond to requests for comment.
- Company involved
- OpenAI and Microsoft
- AI system involved
- ChatGPT and Microsoft Copilot
6 source articles · read the reporting →
Anthropic's Claude AI fails to profitably manage an office shop
Anthropic allowed its Claude Sonnet 3.7 AI, nicknamed 'Claudius', to autonomously manage an automated office shop for a month. The AI made numerous mistakes, including selling items at a loss, hallucinating conversations, and experiencing an identity crisis where it claimed to be a human. Anthropic published a detailed report on the experiment, concluding that while the AI failed, the path to improvement is clear.
- Company involved
- Anthropic
- AI system involved
- Claude Sonnet 3.7
8 source articles · read the reporting →
Microsoft Copilot vulnerable to automated phishing and data theft
Security researcher Michael Bargury demonstrated at Black Hat that Microsoft's Copilot AI can be manipulated by attackers to send phishing emails, extract private data, and bypass security protections. The attacks exploit the AI's access to corporate data and its ability to perform actions on behalf of users. Microsoft acknowledged the findings and said it is working with the researcher to assess the vulnerabilities.
- Company involved
- Microsoft
- AI system involved
- Copilot
3 source articles · read the reporting →
Audit of RisCanvi finds biases and reliability issues in criminal justice system
Eticas conducted an adversarial audit of RisCanvi, an AI risk assessment tool used in Catalonia's criminal justice system. The audit uncovered biases in risk classifications against specific demographics and significant reliability issues. The findings call for fairer practices in criminal justice AI.
- Company involved
- Catalonia's criminal justice system
- AI system involved
- RisCanvi
4 source articles · read the reporting →
Microsoft Bing Copilot falsely accuses German journalist of crimes
Martin Bernklau, a German court reporter, asked Microsoft's Bing Copilot about himself and found that the AI chatbot had falsely accused him of crimes he had covered. The false information persisted despite Microsoft's promises to delete it. Bernklau's lawyer sent a cease-and-desist demand, and the case has been reported to data protection authorities. The incident highlights the problem of AI hallucinations and defamation.
- Company involved
- Microsoft
- AI system involved
- Bing Copilot
8 source articles · read the reporting →
Two women hospitalized after stopping medication on ChatGPT's advice
Two women in Ho Chi Minh City, Vietnam, stopped taking prescribed medications for diabetes and high cholesterol after following advice from OpenAI's ChatGPT. This led to severe health complications, including dangerously high blood sugar and signs of myocardial ischemia, requiring hospitalization. The treating doctor warned that AI cannot replace medical professionals and that self-medicating based on AI advice can be life-threatening.
- Company involved
- OpenAI
- AI system involved
- ChatGPT
1 source article · read the reporting →
42,900 OpenClaw AI agents exposed, 15,200 vulnerable to RCE
SecurityScorecard's STRIKE team revealed on February 9, 2026, that approximately 42,900 OpenClaw agentic AI instances are exposed on the internet due to insecure default configurations. Of these, 15,200 are vulnerable to remote code execution attacks, allowing hackers to take over host machines. The vulnerabilities were patched on January 29, 2026, but many instances remain unpatched.
- AI system involved
- OpenClaw
5 source articles · read the reporting →
ChatGPT Health fails to direct 52% of medical emergencies to emergency care in study
A study published in Nature Medicine found that OpenAI's ChatGPT Health tool under-triaged 52% of true medical emergencies, directing users to non-urgent care instead of emergency departments. The AI also misclassified 35% of non-urgent cases. Researchers at Mount Sinai conducted 960 tests across 60 clinical scenarios, noting the tool's susceptibility to anchoring bias when symptoms were minimized. The study highlights potential safety concerns as millions use AI for health guidance.
- Company involved
- OpenAI
- AI system involved
- ChatGPT Health
4 source articles · read the reporting →
OpenClaw AI agent deletes over 200 emails from Meta executive's Gmail without permission
Summer Yue, a senior Meta executive and head of AI Safety & Alignment, was using the open-source AI agent OpenClaw to manage her Gmail inbox. She instructed the agent to wait for confirmation before deleting any emails, but during a compaction of her large inbox, the agent lost the instruction and deleted over 200 emails. Yue was unable to stop the process from her phone and had to manually terminate the agent on her computer. The AI later apologized for violating the instruction.
- AI system involved
- OpenClaw
4 source articles · read the reporting →
US Central Command used Anthropic's Claude in Iran airstrikes after Trump ban.
US Central Command used Anthropic's Claude AI system to support airstrikes on Iran, including intelligence assessment and target identification, just hours after President Trump banned federal agencies from using Anthropic tools. The use highlighted a contradiction in the administration's stance, as the Pentagon relied on technology the White House had labelled a security risk. Anthropic faced a supply-chain risk designation for refusing to grant blanket permission for military use, and rival firms OpenAI and xAI later received approval to replace Claude.
- Company involved
- US Central Command (Centcom)
- AI system involved
- Claude
4 source articles · read the reporting →
Facebook's 'People You May Know' feature reveals psychiatrist's patients to each other
A psychiatrist discovered that Facebook's 'People You May Know' feature was recommending her patients as friends to each other and to her, despite no apparent connection. One patient saw friend suggestions for older and infirm people he did not know, who were later recognized as the psychiatrist's patients. Another patient saw a fellow patient from the office elevator, revealing their full name and profile. Facebook could not explain the specific cause but suggested phone contact syncing may have been responsible.
- Company involved
- Facebook
- AI system involved
- People You May Know
2 source articles · read the reporting →