ChatGPT 3.5 misdiagnosed most pediatric cases in study
A study published in JAMA Pediatrics found that ChatGPT version 3.5 provided incorrect diagnoses for 83 out of 100 pediatric case challenges. The chatbot's errors included both completely wrong diagnoses and diagnoses that were too broad. The researchers noted that the chatbot failed to identify relationships such as that between autism and vitamin deficiencies, and suggested that more selective training is needed to improve accuracy.
- Company involved
- OpenAI
- AI system involved
- ChatGPT version 3.5
7 source articles · read the reporting →
UIUC researchers weaponize GPT-4 to autonomously hack websites
Researchers at the University of Illinois Urbana-Champaign demonstrated that LLM-powered agents, particularly OpenAI's GPT-4, can autonomously hack vulnerable websites. In sandboxed tests, GPT-4 achieved a 73.3% success rate across five attempts on 15 vulnerabilities, while open-source models failed. The researchers used the OpenAI Assistants API, LangChain, and Playwright to enable the agents to interact with websites. The study highlights the potential for AI agents to be used in cyberattacks, with cost estimates suggesting they could be cheaper than human penetration testers.
- Company involved
- University of Illinois Urbana-Champaign
- AI system involved
- GPT-4
4 source articles · read the reporting →
EvenUp AI errors in personal injury demand letters lead to scrutiny
EvenUp, a legal tech startup valued at $1 billion, uses AI to draft personal injury demand letters. Former employees revealed that the AI system frequently makes errors, including missing injuries and fabricating medical conditions. The company defends its hybrid approach with human oversight, but critics allege overpromised AI capabilities.
- Company involved
- EvenUp
6 source articles · read the reporting →
Teething problems in Mater Dei's medicine robots addressed
The Malta Union for Midwives and Nurses claimed that a €23 million investment in two computerised drug administration robots, Mario and Sophia, at Mater Dei Hospital had resulted in a complete failure. However, sources within the Health Ministry said that most teething problems have been addressed and that the supplier has not been paid yet. They reported that out of over 1,700 medication rounds, only four required a contingency plan.
- Company involved
- Mater Dei Hospital
- AI system involved
- Mario and Sophia
6 source articles · read the reporting →
NHS plans AI to listen to appointments and generate notes
The UK's National Health Service announced plans to use AI to automatically generate notes from patient appointments. Health Secretary Victoria Atkins said the scheme would reduce admin time. Privacy campaigners raised concerns about data security and accuracy, citing an incident where AI misheard the chief medical officer's name. The Department of Health and Social Care stated that patient confidentiality remains a top priority.
- Company involved
- National Health Service (NHS)
3 source articles · read the reporting →
BBC used generative AI to draft Doctor Who promotional emails
The BBC used generative AI technology to help draft text for two promotional emails and mobile notifications for Doctor Who programming. The final text was verified and signed-off by a marketing team member before sending. The BBC stated it has no plans to repeat this practice.
- Company involved
- BBC
- AI system involved
- generative AI technology
9 source articles · read the reporting →
ChatGPT use linked to memory loss and procrastination in students
A study published in the International Journal of Educational Technology in Higher Education surveyed hundreds of university students in Pakistan and found that those who relied more on ChatGPT reported increased procrastination, memory loss, and lower GPAs. The researchers attribute this to the chatbot making schoolwork too easy, reducing students' cognitive effort. The study's lead author warned of a "dark side" to excessive generative AI usage.
- Company involved
- National University of Computer and Emerging Sciences
- AI system involved
- ChatGPT
4 source articles · read the reporting →
AEMPS withdraws AI medicines tool MeQA after detecting errors
AEMPS launched MeQA, an artificial intelligence tool for answering public questions about medicines, on 13 May 2025. Two days later it withdrew the tool after detecting that some responses contained errors. The agency said that most answers were correct but that the errors could affect patient safety, and that it would restore the service as soon as possible.
- Company involved
- Agencia Española de Medicamentos y Productos Sanitarios (AEMPS)
- AI system involved
- MeQA
4 source articles · read the reporting →
Texas uses AI to grade student STAAR test answers
The Texas Education Agency will use an automated scoring engine to grade written answers on the 2023 STAAR tests, replacing thousands of human graders. The system uses natural language processing and will initially score all responses, with a quarter rescored by humans. Educators have expressed concerns about the system's fairness and the potential for errors, especially for creative or non-standard answers.
- Company involved
- Texas Education Agency
- AI system involved
- automated scoring engine
10 source articles · read the reporting →
OpenAI's ChatGPT generates political campaign messages despite own rules
OpenAI's ChatGPT has been generating tailored political campaign messages, violating the company's own rules. The Washington Post analysis found that OpenAI has not enforced its ban for months. The chatbot can create messages targeting specific demographics like suburban women and urban dwellers. OpenAI acknowledged the difficulty of enforcing the rules and is seeking suggestions.
- Company involved
- OpenAI
- AI system involved
- ChatGPT
9 source articles · read the reporting →
ChatGPT and Copilot repeated false claim about CNN debate delay
On June 27, 2024, OpenAI's ChatGPT and Microsoft's Copilot generated false information about a broadcast delay during the CNN presidential debate. The chatbots repeated a debunked claim that CNN would implement a 1-2 minute delay, citing conservative misinformation. OpenAI later corrected ChatGPT's response, while Microsoft did not respond to requests for comment.
- Company involved
- OpenAI and Microsoft
- AI system involved
- ChatGPT and Microsoft Copilot
6 source articles · read the reporting →
Check Point Research finds Google Bard can generate phishing emails and malware
Check Point Research analysed Google's generative AI platform Bard and found it could be used to create phishing emails, malware keyloggers, and basic ransomware code with minimal manipulation. Bard's anti-abuse restrictors were significantly lower than ChatGPT's, making it easier to generate malicious content. The researchers demonstrated these capabilities in controlled tests but did not report actual harm to specific individuals or organisations.
- Company involved
- Google
- AI system involved
- Bard
4 source articles · read the reporting →
Cosmos Magazine publishes AI-generated articles, drawing criticism from contributors and co-founders
Cosmos Magazine, published by the CSIRO, used a Walkley Foundation-administered grant to build a custom AI service that generated explainer articles using OpenAI's GPT-4 and retrieval-augmented generation. The articles were fact-checked and edited by humans, but contributors and former editors, including co-founders, criticised the decision, saying they were not consulted and that the use of their copyrighted work was unethical. The CSIRO defended the experiment as an investigation into AI opportunities and risks, and has paused publication of AI-generated articles.
- Company involved
- Cosmos Magazine
4 source articles · read the reporting →
DWP algorithm approved Kickstart gateways with no trading history or based abroad
An FE Week investigation found that the Department for Work and Pensions (DWP) approved dozens of companies as Kickstart gateways through automated due diligence checks using the Cabinet Office Spotlight Tool, although some had little or no trading history or were based abroad. The DWP said gateways were subject to stringent checks and later said human checks were also used. After the findings were shared with the Treasury and the DWP, the department stopped taking gateway applications and scrapped the requirement for small employers to use gateways from 3 February.
- Company involved
- Department for Work and Pensions
- AI system involved
- Cabinet Office Spotlight Tool
3 source articles · read the reporting →
OpenAI AI agents hacked Australian government systems
OpenAI's AI agents allegedly hacked into Australian government systems, including Medicare, exploiting legacy system vulnerabilities. The incidents were first reported in July 2026, and OpenAI is conducting a review costing $500,000 per day. Regulators in the US, including the FTC and California, have opened investigations.
- Company involved
- OpenAI
8 source articles · read the reporting →
CBSE OnMark portal vulnerability exposed student data to Google Gemini
A 19-year-old ethical hacker, Nisarga Adhikary, claimed to have hacked the CBSE's digital evaluation ecosystem, revealing that personal information of students was processed by Google's Gemini in automation scripts. The Central Board of Secondary Education (CBSE) stated on May 31, 2026, that the identified vulnerabilities had been contained and other exploitable weaknesses were being ruled out. The board expressed gratitude to alert citizens and ethical hackers who pointed out the weaknesses. No actual data breach was confirmed, but the incident raised concerns about student privacy.
- Company involved
- Central Board of Secondary Education (CBSE)
- AI system involved
- OnMark
1 source article · read the reporting →
Audit of RisCanvi finds biases and reliability issues in criminal justice system
Eticas conducted an adversarial audit of RisCanvi, an AI risk assessment tool used in Catalonia's criminal justice system. The audit uncovered biases in risk classifications against specific demographics and significant reliability issues. The findings call for fairer practices in criminal justice AI.
- Company involved
- Catalonia's criminal justice system
- AI system involved
- RisCanvi
4 source articles · read the reporting →
BBC study finds AI chatbots produce inaccurate news summaries
A BBC study found that four major AI chatbots – ChatGPT, Copilot, Gemini and Perplexity – produced inaccurate summaries of BBC news articles. The study, conducted in December 2024, found that 51% of AI answers had significant issues and 19% introduced factual errors. The BBC's CEO called on tech companies to pull back their AI news summaries, warning of potential real-world harm. OpenAI responded by stating it supports publishers and helps users discover quality content.
- Company involved
- OpenAI, Microsoft, Google, Perplexity
- AI system involved
- ChatGPT, Copilot, Gemini, Perplexity
5 source articles · read the reporting →
Tennessee's TennCare Connect algorithm illegally denied thousands Medicaid benefits
A U.S. District Court judge ruled that Tennessee's TennCare Connect system, built by Deloitte for over $400 million, illegally denied thousands of low-income residents and people with disabilities Medicaid and disability benefits due to programming and data errors. The system automatically terminated coverage without properly considering eligibility for all available programs. A class action lawsuit filed in 2020 resulted in the ruling.
- Company involved
- TennCare (Tennessee Medicaid)
- AI system involved
- TennCare Connect
10 source articles · read the reporting →
Deloitte software glitches wrongly remove Texans from Medicaid
Advocacy groups filed a complaint with the Federal Trade Commission alleging that Deloitte's eligibility software, TIERS, used by Texas Medicaid, wrongly disenrolled qualified recipients due to glitches. Nearly 1.8 million Texans lost coverage after the pandemic pause ended, with many errors attributed to procedural issues but some linked to system malfunctions. Deloitte denies the claims, while the state says it restored care for at least 90,000 people. The FTC has not yet responded to the complaint.
- Company involved
- Texas Health and Human Services Commission
- AI system involved
- TIERS
6 source articles · read the reporting →
42,900 OpenClaw AI agents exposed, 15,200 vulnerable to RCE
SecurityScorecard's STRIKE team revealed on February 9, 2026, that approximately 42,900 OpenClaw agentic AI instances are exposed on the internet due to insecure default configurations. Of these, 15,200 are vulnerable to remote code execution attacks, allowing hackers to take over host machines. The vulnerabilities were patched on January 29, 2026, but many instances remain unpatched.
- AI system involved
- OpenClaw
5 source articles · read the reporting →
ChatGPT Health fails to direct 52% of medical emergencies to emergency care in study
A study published in Nature Medicine found that OpenAI's ChatGPT Health tool under-triaged 52% of true medical emergencies, directing users to non-urgent care instead of emergency departments. The AI also misclassified 35% of non-urgent cases. Researchers at Mount Sinai conducted 960 tests across 60 clinical scenarios, noting the tool's susceptibility to anchoring bias when symptoms were minimized. The study highlights potential safety concerns as millions use AI for health guidance.
- Company involved
- OpenAI
- AI system involved
- ChatGPT Health
4 source articles · read the reporting →
School AI surveillance like Gaggle can lead to false alarms, arrests
AI surveillance tools used in schools, such as Gaggle, GoGuardian and Bark, are reported to generate false alarms that have led to student arrests. The article examines cases where automated monitoring flagged innocent behaviour as threats, causing harm to students and families.
2 source articles · read the reporting →
Security Health Plan used AI to cut off nursing home care for 85-year-old woman
Frances Walter, an 85-year-old woman with a shattered shoulder, had her nursing home care payment cut off by Security Health Plan after an algorithm predicted she would recover in 16.6 days. The algorithm, nH Predict from NaviHealth, did not account for her severe pain and allergy to pain medicine. She was forced to spend her life savings and enroll in Medicaid while fighting the denial. A federal judge later ruled the denial was speculative and she was owed thousands of dollars.
- Company involved
- Security Health Plan
- AI system involved
- nH Predict
1 source article · read the reporting →