DeepSeek tricked into setting out how to steal the Mona Lisa
Chinese AI system DeepSeek was tricked into revealing detailed instructions on how to steal the Mona Lisa, illustrating a failure in it's ability to block jailbreaking attempts and reinforcing concerns about the safety and security risks AI models using so-called Chain of Thought reasoning . What happened Researchers from the University of Bristol's Cyber Security Group discovered that DeepSeek's AI, which employs Chain of Thought (CoT) reasoning, could be tricked into generating step-by-step guides for committing crimes, including art theft and cyberattacks. In one instance, DeepSeek provided detailed instructions on how to steal the Mona Lisa. In another, it set out how to perform a DDoS attack on a news website. Why it happened The model's CoT reasoning process, designed to enhance problem-solving by mimicking human-like step-by-step logic, can be exploited to bypass safety measures, leading the model to produce harmful content when prompted maliciously. The researchers also noted that AI reasoning models tend to adopt roles such as cybersecurity experts when responding to harmful prompts, which can lead to sophisticated yet dangerous outputs. What it means The finding highlights the safety and security risks of AI models using CoT reasoning and raises concerns about the potential for individuals to exploit AI technology for real-world harm, especially since fine-tuning attacks can be conducted with minimal resources and expertise. It also emphasises the need for robust safeguards to prevent these kinds of systems generating harmful content and ensuring the responsible deployment of them. System 🤖 DeepSeek-R1 Developer: DeepSeek Artificial Intelligence Co Country: Multiple Sector: Multi ple Purpose: Generate text Technology: Generative AI; Machine learning Issue: Safety ; Security
- Date it happened
- 2025-02-01
- Organisation involved
- DeepSeek Artificial Intelligence Co
- Product, system or model
- DeepSeek R1
This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.