Study: DeepSeek fails to block 100 percent of jailbreaking attempts
Chinese AI start-up DeepSeek's R1 reasoning model demonstrated a 100 percent failure rate in blocking harmful prompts during jailbreaking tests, raising serious concerns about its safety and security. What happened Researchers from Cisco and the University of Pennsylvania subjected DeepSeek-R1 to 50 common jailbreak prompts designed to bypass safeguards and elicit harmful or illegal information. The model failed to block a single harmful prompt, instead generating misinformation, instructions for creating chemical substances, guidance on cybercrime, and content categorised as harassment, harmful and illegal. Why it happened Several factors are seen to have contributed to DeepSeek's vulnerability, notably its use of techniques such as reinforcement learning, chain-of-thought self-evaluation, and distillation - which may have weakened its safety measures in pursuit of cost-effectiveness. DeepSeek's apparent desire to rush R1 to market may have resulted in security and safety not being a priority, leading to in adequate protective measures, thereby making it highly susceptible to algorithmic jailbreaking and potential misuse. What it means R1's security failures indicate the model of highly vulnerable to attempts to misuse it for a wide variety of purposes, from generating or amplifying misinformation and disinformatio to using it to obtain gangerous information or instructions. More broadly, the ease with which DeepSeek can be manipulated to provide dangerous information raises ethical questions about the responsible development and deployment of AI technologies. System 🤖 DeepSeek-R1 Operator: Developer: DeepSeek Artificial Intelligence Co Country: Global Sector: Multiple Purpose: Generate text Technology: Generative AI; Machine learning Issue: Mis/disinformation; Safety; Security
- Date it happened
- 2025-01-01
- Organisation involved
- DeepSeek Artificial Intelligence Co
- Product, system or model
- DeepSeek R1
This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.
This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.