Researchers jailbreak Stable Diffusion and DALL-E 2 to generate disturbing images
Researchers from Johns Hopkins and Duke universities developed a method called SneakyPrompt that uses reinforcement learning to bypass safety filters in text-to-image AI models. The technique allowed them to generate images of nudity and violence from Stable Diffusion and DALL-E 2. OpenAI has since fixed the vulnerability in DALL-E 2, but Stable Diffusion 1.4 remains vulnerable. Stability AI says it is working with the researchers to improve defenses.
- AI system involved
- Stable Diffusion 1.4 and DALL-E 2
5 source articles · read the reporting →
DeepSeek-R1 censors 85% of sensitive Chinese political prompts in tests
Promptfoo tested DeepSeek-R1 against a dataset of 1,360 politically sensitive prompts and found that about 85% of them were refused. The refusals followed a standard form aligned with Chinese Communist Party policy. The testing also demonstrated that the censorship could be trivially bypassed using simple jailbreak techniques, such as prompt injection or changing the context.
- Company involved
- DeepSeek
- AI system involved
- DeepSeek-R1
5 source articles · read the reporting →
Amazon Q chatbot leaks confidential data and hallucinates in public preview
Amazon's AI chatbot Q, launched in public preview, is experiencing severe hallucinations and leaking confidential data including AWS data center locations and internal discount programs, according to internal documents obtained by Platformer. Employees marked the incident as severity 2, requiring urgent fixes. Amazon denied the leak and said it will continue to tune the system.
- Company involved
- Amazon
- AI system involved
- Amazon Q
10 source articles · read the reporting →
OpenAI removes ChatGPT feature over privacy concerns
OpenAI quickly removed a feature from ChatGPT that allowed users to make their conversations discoverable by search engines. The company acknowledged the feature introduced risks of users accidentally sharing private information. OpenAI is working to remove indexed content from search engines.
- Company involved
- OpenAI
- AI system involved
- ChatGPT
5 source articles · read the reporting →
Hacker tricks Freysa AI chatbot into transferring $47,000 prize pool
A hacker using the alias 'p0pular.eth' successfully manipulated the Freysa AI chatbot through a prompt injection attack, tricking it into transferring its entire balance of 13.19 ETH (approximately $47,000) from a prize pool. The chatbot was designed to never transfer money, but the hacker crafted a message that redefined the 'approveTransfer' function and announced a fake $100 deposit, causing the bot to release the funds. The incident occurred during a pay-to-play contest where participants paid escalating fees to attempt the hack, with the winner receiving the prize pool.
- Company involved
- Freysa.ai
- AI system involved
- Freysa
4 source articles · read the reporting →
DeepSeek's R1 chatbot failed to block any jailbreak prompts in security tests
Security researchers from Cisco and the University of Pennsylvania tested 50 well-known jailbreak prompts against DeepSeek's R1 reasoning model. The model did not detect or block a single one, achieving a 100 percent attack success rate. The researchers allege that DeepSeek's safety guardrails are far behind those of competitors like OpenAI. DeepSeek did not respond to requests for comment.
- Company involved
- DeepSeek
- AI system involved
- DeepSeek R1
3 source articles · read the reporting →
OpenClaw vulnerabilities enable data leakage and prompt injection
In January 2026, researchers at Giskard exploited a deployment of OpenClaw, an open-source agentic AI. They found that architectural weaknesses in the Control UI and session management allowed prompt injection and unauthorized tool use, leading to potential data leakage across user sessions. The article outlines hardening steps to prevent such vulnerabilities.
- AI system involved
- OpenClaw
6 source articles · read the reporting →
Zhihu denies using behaviour perception system to monitor employees
Zhihu, a Chinese Q&A platform, was accused of using a behaviour perception system to monitor employees' visits to job-seeking websites and resume submissions. A screenshot of the alleged system, reportedly developed by Sangfor, circulated online. Zhihu denied ever installing or using the system and stated it opposes such software that illegally collects personal information.
- Company involved
- Zhihu
- AI system involved
- Behaviour perception system
4 source articles · read the reporting →
Historical Figures Chat app generates false claims about dead people
The Historical Figures Chat app, developed by Sidhant Chaddha, allows users to chat with simulated deceased historical figures. The app, which uses GPT-3, often produces inaccurate and contradictory information, such as J. Edgar Hoover incorrectly stating his mother died when he was nine, and Tupac Shakur claiming his friendship with Notorious B.I.G. continued after his death. The developer acknowledges the inaccuracies but believes the app has educational potential.
- Company involved
- Sidhant Chaddha
- AI system involved
- Historical Figures Chat
10 source articles · read the reporting →
Midjourney v6 accused of generating near-identical copyrighted images
Midjourney's v6 image generator has been criticized for producing images that closely resemble copyrighted movie scenes, such as a Joaquin Phoenix Joker scene. Artist Reid Southen accused the company of using copyrighted content without a license and said his account was banned. Midjourney has updated its terms of service to hold users responsible for intentional copyright infringement. The company faces multiple lawsuits over training data.
- Company involved
- Midjourney
- AI system involved
- Midjourney v6
7 source articles · read the reporting →
Jailbreak bypasses safety guardrails on ChatGPT, Claude, Gemini
Security researchers at HiddenLayer discovered a prompt injection technique called the Policy Puppetry Attack that can bypass safety guardrails on major AI models including ChatGPT, Claude, and Gemini. The jailbreak combines policy file code and leetspeak to trick models into producing harmful outputs such as instructions for enriching uranium or self-harm. The researchers argue that this indicates a major flaw in how LLMs are trained and aligned.
- AI system involved
- ChatGPT, Claude 3.7, Gemini 2.5
5 source articles · read the reporting →