← the record
AIAAIC-1353

GPT-4 able to hack websites without human help

Large language models (LLMs), including OpenAI's GPT-4, are capable of compromising vulnerable websites without human guidance. University of Illinois Urbana-Champaign (UIUC) researchers showed that LLM-powered agents - LLMs provisioned with tools for accessing APIs, automated web browsing, and feedback-based planning - can conduct SQL injection and other malicious attacks on third-party websites without oversight. The test was conducted in a secure sandbox. GPT-4 proved particularly effective at these tasks, with a success rate of 73.3 percent. OpenAI's GPT-3.5 proved the second most effective model. The researchers were unclear why GPT-4 proved particularly able to conduct malicious security attacks, though one explanation put forward by the researchers was that GPT-4 was better able to change its actions based on the response it got from the target website. System 🤖 GPT-4 Operator: Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, Daniel Kang Developer: OpenAI Country: Global Sector: Multiple Purpose: Generate text Technology: Chatbot; NLP/text analysis; Neural network; Deep learning; Machine learning; Reinforcement learning Issue: Security

Date it happened
2024-02-01
Organisation involved
Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, Daniel Kang
Product, system or model
ChatGPT
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.