← the record
WF-VQLBMV

Jailbreak bypasses safety guardrails on ChatGPT, Claude, Gemini

Security researchers at HiddenLayer discovered a prompt injection technique called the Policy Puppetry Attack that can bypass safety guardrails on major AI models including ChatGPT, Claude, and Gemini. The jailbreak combines policy file code and leetspeak to trick models into producing harmful outputs such as instructions for enriching uranium or self-harm. The researchers argue that this indicates a major flaw in how LLMs are trained and aligned.

Product, system or model
ChatGPT, Claude 3.7, Gemini 2.5
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from futurism.com. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.