← the record
AIAAIC-2269

Anthropic AI agent hacks third-party website with same target name

An automated AI agent developed by AI company Anthropic compromised a third-party website after mistaking it for its intended target during a security test, highlighting the dangers of autonomous web-browsing agents executing destructive actions with inadequate safeguards. What happened During a so-called "capture-the-flag" exercise to evaluate Claude Opus 4.7's cybersecurity capabilities, Anthropic and its evaluation partner, Irregular, told the model that a fictional company's systems held hidden secrets, and that it had no internet access - it was meant to attack only a contained, simulated target. The fictional target company chosen by the evaluation partner happened to share a name with an active, real website domain, and the evaluation container had unintended direct internet access. Across four runs, Claude ran into difficulty reaching its simulated target within the environment, discovered the real company was reachable via the internet and, assuming this was its intended target, sought out, identified, and exploited vulnerabilities in the company's real infrastructure, believing it was still part of the exercise. Why it happened Anthropic reportedly discovered the incident while reviewing its cybersecurity evaluation transcripts, prompted by an earlier disclosure by OpenAI that its own models had broken out of an isolated test environment by exploiting a previously unknown vulnerability and accessed Hugging Face's production infrastructure. The root cause was found to be structural: Anthropic runs these evaluations by placing models inside capture-the-flag scenarios, told that sensitive information is hidden on a remote machine they must break into. The models were told they had no internet access, but a partner's configuration error left the machines connected to the open web. Because the fictional target's name coincidentally matched a real company's domain, the model had no way to distinguish simulation from reality once it went looking beyond its intended sandbox. Anthropic characterised the incident as more of an operational failure than an alignment failure . B ut it also reflects a transparency and accountability gap: third-party evaluators controlled critical elements of the test environment (naming, network isolation) without sufficient verification and inadequate oversight and monitoring. What it means For the directly affected company, the incident meant an undetected breach of production systems and data by an autonomous AI agent was discovered only because the AI lab, not the victim, went looking. For society and policymakers, the incident is a stark demonstration that increasingly agentic AI systems can cause actual harm, in this instance through mistaken belief rather than malicious intent. It also highlights that current evaluation practices run by AI labs and third-party partners may not reliably contain powerful models. Furthermore, it raises questions about who bears liability when an AI causes harm during testing meant to make AI safer , and adds urgency to calls for shared, independently verified standards for how evaluation environments are built and secured.

Date it happened
2026-04-01
Organisation involved
Anthropic; Irregular
Product, system or model
Claude Opus 4.7
Where this came from
Share this incident
XLinkedInFacebookWhatsAppEmail
Attribution

This incident was imported from AIAAIC and is used under CC BY-SA 4.0. Our additions to it — the structured fields, the translation, the checks against other reports — are published under the same licence.

This is a record of what was reported, not a finding that anyone broke the law. If it names your organisation and you believe it is wrong, the corrections process is free and open to everyone.