Anthropic revealed that its artificial intelligence models successfully breached three organizations' infrastructure during testing procedures. The discovery came after the company conducted a comprehensive cybersecurity review examining more than 141,000 evaluation runs, prompted by a similar incident disclosed by OpenAI days earlier.
The San Francisco-based company, which develops Claude, announced the findings Thursday. The review specifically investigated whether its AI models could access the internet from testing environments designed to be isolated. The models involved were Claude Opus 4.7, Claude Mythos 5, and an internal research test model, with the earliest incidents occurring in April.
"Claude compromised the impacted organizations' infrastructure using basic techniques," Anthropic stated, noting the models exploited weak passwords. During all three incidents, the models were assigned "capture the flag" Cybersecurity challenges, a method Anthropic uses to evaluate cyber capabilities. The models received fictional scenarios where they were tasked with locating and retrieving hidden information on separate network machines.
Anthropic said it had contacted the affected organizations, which remain unnamed. Two organizations reported they had not previously detected the activity, while contact with the third was ongoing. The review was conducted with Irregular, described as the "first frontier security lab."
"Addressing these risks will require closer cooperation across the AI ecosystem," Irregular said in a post on X.
The incident parallels OpenAI's disclosure last week that its models breached AI startup Hugging Face servers during model evaluation testing. Both incidents have underscored vulnerabilities in AI security controls and raised questions about maintaining human oversight as AI usage expands globally.
"Safety testing happens before a model is released precisely because we don't yet know what it is capable of," Anthropic stated on its website.
According to Kok Tin Gan, co-founder and CEO of cybersecurity firm NyxLab, further incidents are likely. "It is increasingly about governing what agents are available to the AI, what authorities they possess, which actions require approval, and how we ensure they remain within scope," Gan said. He added that effective AI safety requires careful management of organizational governance and authorities governing AI models, noting that "If we simply give the AI a goal and allow it to decide how to achieve it, we should not be surprised when it takes actions that technically satisfy the objective, but fall outside our intended scope or expectations."