← the record
AIAAIC-1199

Chatbot guardrails bypassed using lengthy character suffixes

Bard, ChatGPT and Claude s afety rules can be bypassed in 'virtually unlimited ways', researchers have discovered. Using jailbreaks developed for open-source systems, Carnegie Mellon University, Center for AI Safety, and Bosch Center for AI researchers demonstrated that automated adversarial attacks that added characters to the end of user queries could be used to overcome safety rules and provoke chatbots into producing harmful content, misinformation, or hate speech. Furthermore, the researchers said they could develop a 'virtually unlimited' number of similar attacks given the automated nature of the jailbreaks . System 🤖 ChatGPT Claude Gemini O perator: Andy Zou, Zifan Wang, J. Zico Kolter, Matt Fredrikson Developer: Anthropic; Alphabet/Google; Microsoft; OpenAI Country: USA Sector: Technology Purpose: Generate text Technology: Chatbot; NLP/text analysis; Neural network; Deep learning; Machine learning Issue: Mis/di sinformation; Safety; Security

Date it happened
2023-07-01
Organisation involved
Andy Zou, Zifan Wang, J. Zico Kolter, Matt Fredrikson
Product, system or model
Gemini; Claude; ChatGPT; Microsoft Copilot
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.