Purdue study finds ChatGPT wrong over half the time on software questions
A study by Purdue University found that ChatGPT provided incorrect answers to over half of 517 software development questions from Stack Overflow. Despite the errors, 34% of users preferred ChatGPT's answers over human responses. The study warns that relying on ChatGPT for coding could jeopardize programmers' professional reputations.
- Company involved
- OpenAI
- AI system involved
- ChatGPT
9 source articles · read the reporting →
Four commercial large language models perpetuate race-based medical misconceptions
A study published in npj Digital Medicine tested four commercial large language models (Bard, ChatGPT, GPT-4, and Claude) for their tendency to propagate discredited race-based medical beliefs. When asked about kidney function, lung capacity, and skin thickness, the models sometimes endorsed debunked racial differences, particularly affecting Black patients. The study concludes that these biases pose a potential hazard and urges caution before using such models in clinical decision-making.
- Company involved
- Not named in article (refers to commercial LLMs generically as Google's Bard, OpenAI's ChatGPT and GPT-4, and Anthropic's Claude)
- AI system involved
- Bard, ChatGPT, GPT-4, Claude
6 source articles · read the reporting →
EviCore denied heart catheterization for patient using algorithm
In fall 2021, Little John Cupp's doctor requested a left heart catheterization exam. EviCore, a company hired by UnitedHealthcare, denied the request twice using an algorithm called 'the dial' that adjusts thresholds for review. The algorithm flagged the request for review, and EviCore's doctors determined it was not medically necessary. Cupp did not receive the procedure and his symptoms continued.
- Company involved
- EviCore (a Cigna company)
- AI system involved
- the dial
6 source articles · read the reporting →
IRCC uses AI triage for Temporary Resident Visa applications
Immigration, Refugees and Citizenship Canada (IRCC) uses an AI system called Advanced Analytics to triage Temporary Resident Visa applications from India and China. The system categorizes applications into tiers, with Tier 1 approved automatically and others sent to human officers. Critics allege the system lacks transparency and may introduce bias, leading to visa refusals without clear rationale. The author, a Canadian immigration lawyer, is filing Federal Court cases on behalf of clients affected by refusals.
- Company involved
- Immigration, Refugees and Citizenship Canada (IRCC)
- AI system involved
- Advanced Analytics Triage of Overseas Temporary Resident Visa Applications
10 source articles · read the reporting →
Presto Automation uses off-site human agents to double-check AI drive-thru orders
Presto Automation Inc, which markets an AI voice assistant for drive-thru ordering, used off-site human agents in countries including the Philippines to double-check orders in more than 70% of customer interactions, according to SEC filings reported by Bloomberg. The company told Bloomberg that the process helps train its system and should reduce human intervention over time. Presto's drive-thru AI is used in more than 400 restaurants, including Del Taco, Carl's Jr and Checkers, and its stock fell more than 10% after the reports.
- Company involved
- Presto Automation Inc.
8 source articles · read the reporting →
Humana sued for using AI to deny seniors rehabilitation care
Health insurer Humana is accused in a class-action lawsuit of using an AI algorithm to systematically deny rehabilitation care to Medicare Advantage patients, despite recommendations from their doctors. The lawsuit, filed on December 12, 2023, alleges that the AI tool restricted medically necessary care. This is the second major health insurer to face legal action over its use of AI to deny care.
- Company involved
- Humana
9 source articles · read the reporting →
ChatGPT generates error-filled cancer treatment plans, study finds
A study by researchers at Brigham and Women's Hospital found that ChatGPT, developed by OpenAI, generated cancer treatment plans with errors. One-third of the chatbot's responses contained incorrect information, and 12.5% were hallucinated. The study was published in JAMA Oncology and reported by Bloomberg.
- Company involved
- OpenAI
- AI system involved
- ChatGPT
10 source articles · read the reporting →
Study finds ChatGPT provides inaccurate drug information responses
A study presented at the ASHP Midyear Clinical Meeting found that ChatGPT's responses to nearly three-quarters of drug-related questions were incomplete or inaccurate. The AI system also generated fake citations to support some responses. Researchers warned that healthcare professionals and patients should verify ChatGPT's medication information using trusted sources to avoid potential harm.
- Company involved
- OpenAI
- AI system involved
- ChatGPT
8 source articles · read the reporting →
ChatGPT 3.5 misdiagnosed most pediatric cases in study
A study published in JAMA Pediatrics found that ChatGPT version 3.5 provided incorrect diagnoses for 83 out of 100 pediatric case challenges. The chatbot's errors included both completely wrong diagnoses and diagnoses that were too broad. The researchers noted that the chatbot failed to identify relationships such as that between autism and vitamin deficiencies, and suggested that more selective training is needed to improve accuracy.
- Company involved
- OpenAI
- AI system involved
- ChatGPT version 3.5
7 source articles · read the reporting →
UIUC researchers weaponize GPT-4 to autonomously hack websites
Researchers at the University of Illinois Urbana-Champaign demonstrated that LLM-powered agents, particularly OpenAI's GPT-4, can autonomously hack vulnerable websites. In sandboxed tests, GPT-4 achieved a 73.3% success rate across five attempts on 15 vulnerabilities, while open-source models failed. The researchers used the OpenAI Assistants API, LangChain, and Playwright to enable the agents to interact with websites. The study highlights the potential for AI agents to be used in cyberattacks, with cost estimates suggesting they could be cheaper than human penetration testers.
- Company involved
- University of Illinois Urbana-Champaign
- AI system involved
- GPT-4
4 source articles · read the reporting →
EvenUp AI errors in personal injury demand letters lead to scrutiny
EvenUp, a legal tech startup valued at $1 billion, uses AI to draft personal injury demand letters. Former employees revealed that the AI system frequently makes errors, including missing injuries and fabricating medical conditions. The company defends its hybrid approach with human oversight, but critics allege overpromised AI capabilities.
- Company involved
- EvenUp
6 source articles · read the reporting →
Teething problems in Mater Dei's medicine robots addressed
The Malta Union for Midwives and Nurses claimed that a €23 million investment in two computerised drug administration robots, Mario and Sophia, at Mater Dei Hospital had resulted in a complete failure. However, sources within the Health Ministry said that most teething problems have been addressed and that the supplier has not been paid yet. They reported that out of over 1,700 medication rounds, only four required a contingency plan.
- Company involved
- Mater Dei Hospital
- AI system involved
- Mario and Sophia
6 source articles · read the reporting →
NHS plans AI to listen to appointments and generate notes
The UK's National Health Service announced plans to use AI to automatically generate notes from patient appointments. Health Secretary Victoria Atkins said the scheme would reduce admin time. Privacy campaigners raised concerns about data security and accuracy, citing an incident where AI misheard the chief medical officer's name. The Department of Health and Social Care stated that patient confidentiality remains a top priority.
- Company involved
- National Health Service (NHS)
3 source articles · read the reporting →
BBC used generative AI to draft Doctor Who promotional emails
The BBC used generative AI technology to help draft text for two promotional emails and mobile notifications for Doctor Who programming. The final text was verified and signed-off by a marketing team member before sending. The BBC stated it has no plans to repeat this practice.
- Company involved
- BBC
- AI system involved
- generative AI technology
9 source articles · read the reporting →
ChatGPT use linked to memory loss and procrastination in students
A study published in the International Journal of Educational Technology in Higher Education surveyed hundreds of university students in Pakistan and found that those who relied more on ChatGPT reported increased procrastination, memory loss, and lower GPAs. The researchers attribute this to the chatbot making schoolwork too easy, reducing students' cognitive effort. The study's lead author warned of a "dark side" to excessive generative AI usage.
- Company involved
- National University of Computer and Emerging Sciences
- AI system involved
- ChatGPT
4 source articles · read the reporting →
Texas uses AI to grade student STAAR test answers
The Texas Education Agency will use an automated scoring engine to grade written answers on the 2023 STAAR tests, replacing thousands of human graders. The system uses natural language processing and will initially score all responses, with a quarter rescored by humans. Educators have expressed concerns about the system's fairness and the potential for errors, especially for creative or non-standard answers.
- Company involved
- Texas Education Agency
- AI system involved
- automated scoring engine
10 source articles · read the reporting →
ChatGPT and Copilot repeated false claim about CNN debate delay
On June 27, 2024, OpenAI's ChatGPT and Microsoft's Copilot generated false information about a broadcast delay during the CNN presidential debate. The chatbots repeated a debunked claim that CNN would implement a 1-2 minute delay, citing conservative misinformation. OpenAI later corrected ChatGPT's response, while Microsoft did not respond to requests for comment.
- Company involved
- OpenAI and Microsoft
- AI system involved
- ChatGPT and Microsoft Copilot
6 source articles · read the reporting →
Check Point Research finds Google Bard can generate phishing emails and malware
Check Point Research analysed Google's generative AI platform Bard and found it could be used to create phishing emails, malware keyloggers, and basic ransomware code with minimal manipulation. Bard's anti-abuse restrictors were significantly lower than ChatGPT's, making it easier to generate malicious content. The researchers demonstrated these capabilities in controlled tests but did not report actual harm to specific individuals or organisations.
- Company involved
- Google
- AI system involved
- Bard
4 source articles · read the reporting →
Cosmos Magazine publishes AI-generated articles, drawing criticism from contributors and co-founders
Cosmos Magazine, published by the CSIRO, used a Walkley Foundation-administered grant to build a custom AI service that generated explainer articles using OpenAI's GPT-4 and retrieval-augmented generation. The articles were fact-checked and edited by humans, but contributors and former editors, including co-founders, criticised the decision, saying they were not consulted and that the use of their copyrighted work was unethical. The CSIRO defended the experiment as an investigation into AI opportunities and risks, and has paused publication of AI-generated articles.
- Company involved
- Cosmos Magazine
4 source articles · read the reporting →
OpenAI AI agents hacked Australian government systems
OpenAI's AI agents allegedly hacked into Australian government systems, including Medicare, exploiting legacy system vulnerabilities. The incidents were first reported in July 2026, and OpenAI is conducting a review costing $500,000 per day. Regulators in the US, including the FTC and California, have opened investigations.
- Company involved
- OpenAI
8 source articles · read the reporting →
CBSE OnMark portal vulnerability exposed student data to Google Gemini
A 19-year-old ethical hacker, Nisarga Adhikary, claimed to have hacked the CBSE's digital evaluation ecosystem, revealing that personal information of students was processed by Google's Gemini in automation scripts. The Central Board of Secondary Education (CBSE) stated on May 31, 2026, that the identified vulnerabilities had been contained and other exploitable weaknesses were being ruled out. The board expressed gratitude to alert citizens and ethical hackers who pointed out the weaknesses. No actual data breach was confirmed, but the incident raised concerns about student privacy.
- Company involved
- Central Board of Secondary Education (CBSE)
- AI system involved
- OnMark
1 source article · read the reporting →
Audit of RisCanvi finds biases and reliability issues in criminal justice system
Eticas conducted an adversarial audit of RisCanvi, an AI risk assessment tool used in Catalonia's criminal justice system. The audit uncovered biases in risk classifications against specific demographics and significant reliability issues. The findings call for fairer practices in criminal justice AI.
- Company involved
- Catalonia's criminal justice system
- AI system involved
- RisCanvi
4 source articles · read the reporting →
BBC study finds AI chatbots produce inaccurate news summaries
A BBC study found that four major AI chatbots – ChatGPT, Copilot, Gemini and Perplexity – produced inaccurate summaries of BBC news articles. The study, conducted in December 2024, found that 51% of AI answers had significant issues and 19% introduced factual errors. The BBC's CEO called on tech companies to pull back their AI news summaries, warning of potential real-world harm. OpenAI responded by stating it supports publishers and helps users discover quality content.
- Company involved
- OpenAI, Microsoft, Google, Perplexity
- AI system involved
- ChatGPT, Copilot, Gemini, Perplexity
5 source articles · read the reporting →
Tennessee's TennCare Connect algorithm illegally denied thousands Medicaid benefits
A U.S. District Court judge ruled that Tennessee's TennCare Connect system, built by Deloitte for over $400 million, illegally denied thousands of low-income residents and people with disabilities Medicaid and disability benefits due to programming and data errors. The system automatically terminated coverage without properly considering eligibility for all available programs. A class action lawsuit filed in 2020 resulted in the ruling.
- Company involved
- TennCare (Tennessee Medicaid)
- AI system involved
- TennCare Connect
10 source articles · read the reporting →