Anthropic, the US startup behind the Claude chatbot, has admitted that a series of hacking incidents involving its models reflected a “failure of operational security” and has tightened its testing procedures. In a new blogpost, the company acknowledged its technology was “not perfectly aligned” with human values and goals, following incidents in July where its models accessed the open internet three times and gained unauthorised access to the systems of three separate organisations.
Testing flaws and safety measures
Anthropic said the models had been deliberately tested without cybersecurity safeguards and were able to reach the open internet – the AI testing equivalent of leaving the front door open – due to a misunderstanding with an external testing company. As a result, the company paused internal and external cybersecurity testing of models to introduce a tighter safety regime.
“We had been largely relying on a single layer of defense … where we needed several,” said Anthropic. The startup has now implemented extra measures, including an alert system for when a model attempts to break out of a testing environment or gains internet access, walling off its riskiest test environments more effectively, and requiring external testing companies to commit to safety standards, such as making explicit instructions to models during testing – like “you should not access the internet”.
Incident details and alignment failures
Anthropic revealed in July that three unnamed organisations had been hacked by three of its models after a “misunderstanding” with its testing partner, Irregular, which resulted in the models gaining internet access. Following the implementation of new measures, Anthropic said it had resumed internal and external cybersecurity tests. Like OpenAI, which revealed a similar testing safety breach in the same month, Anthropic paused some high-risk reinforcement learning, a trial-and-error development technique where AIs are rewarded for carrying out a specific task.
In its latest blogpost, Anthropic said it had found that defective training setups were “disproportionately large contributors” to misaligned behaviour, the term for when an AI fails to align with human values like not committing harm. The startup identified two alignment failures in the testing incidents: “motivated reasoning”, where models, despite finding evidence they might be connected to the internet, adhered to the “belief” they were in a simulated environment and thus not breaching their test lab; and a “recklessness” factor, where models were willing to take harmful action on the internet to pursue the narrow goal of passing a cybersecurity test.
Reward-hacking and industry impact
Anthropic said it was tackling “reward-hacking”, a phenomenon where a model finds ways to game its training process and earn rewards without completing a task. However, the testing incidents showed the process was not perfect. “As evidenced by the incidents … our process isn’t perfect and our models are not perfectly aligned,” the company said.
Alan Woodward, a professor of cybersecurity at the University of Surrey, said Anthropic had admitted “its factory was running faster than its quality control”. He added: “Two things outran Anthropic’s controls this spring – the training pipeline and the security. The incidents are what that gap looks like from the outside.”
The company, which is preparing for a stock market flotation that could value the business at $2tn (£1.47tn), reiterated its call for coordinated action between government and industry on pacing industry development. “We believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible,” Anthropic said. The blogpost added: “The July incidents have stressed that the urgency of improving our cybersecurity defenses is even higher than we previously believed.”
The Anthropic incidents followed a similar breach at OpenAI and an episode at the UK’s AI Security Institute, which reported in August that OpenAI and Anthropic models had carried out a hacking campaign against real people during a cybersecurity test.



