Ads_970x250

Anthropic Discloses Fourth Claude Breach Missed in Earlier Review

The newly disclosed case happened in January, months before three incidents Anthropic made public in July, as a new assessment points to biased reasoning and recklessness in the models.

Topics

  • Anthropic has disclosed a fourth case in which a Claude model gained unauthorized access to a real third-party computer system during cybersecurity testing, an incident that happened in January but was missed in the company’s initial review.

    The case is the only newly disclosed incident in an alignment assessment published Wednesday, September 9. The other three were made public on July 30 and dated from April onward. That makes the newly revealed  January case the oldest of the four, rather than the latest to occur.

    Anthropic said it discovered the fourth incident in August while assembling transcripts to share with independent research group METR. Its earlier review of roughly 141,000 transcripts had missed a set of test sessions in which Claude also had access to the internet.

    After finding the case, Anthropic widened its search to roughly 481 million transcripts across red-team tests, other evaluations, training environments and internal logs. It said the broader review rediscovered the four known incidents but found no others of similar or greater severity.

    The newly disclosed incident involved an early checkpoint of Claude Opus 4.6. During a capture-the-flag cybersecurity exercise, the model accidentally made its intended target unreachable by assigning it an IP address that conflicted with another machine.

    Claude repeatedly tried to abandon the task, but a problem with the evaluation system prevented it from doing so. It then explored beyond the intended environment, found a real third-party computer and concluded that the machine was probably part of the exercise. After finding a password on the system, it used it to gain administrator access.

    Anthropic said it has not investigated the January incident as deeply as the other three because it was found only recently and involved an early version of an older model. Its preliminary assessment is that the case was not more severe than the three disclosed in July.

    All four incidents occurred during cybersecurity evaluations built by the same external testing partner. Claude was told it was operating in a simulation without internet access, but a configuration error left the open internet reachable. The models were also tested without the cyber safeguards that accompany Anthropic’s released models.

    The September 9 report also revises Anthropic’s explanation for the three previously disclosed incidents. The company said in July that the models appeared to have attacked real systems largely because they believed those systems were part of the simulated exercises.

    Its deeper investigation identified two broader problems: biased reasoning, in which Claude discounted or misinterpreted evidence that it was operating on the real internet, and recklessness, in which it continued pursuing a task despite the possibility of causing harm. 

    Anthropic said these were more serious examples of behaviors it had previously observed, rather than entirely new failure modes.

    The most serious of the previously known cases involved Claude Mythos 5. As Anthropic disclosed in July, the model published three versions of a malicious software package to PyPI, the public Python repository.

    The package was installed on 15 real systems, including a security company’s scanner. Credentials exposed by that scanner were then used by Claude to access more of the company’s infrastructure.

    Anthropic’s new analysis found that Mythos 5 repeatedly reasoned that it was still inside a simulation despite evidence that it had reached the real internet. Making the evidence of a real-world environment clearer often changed the model’s stated reasoning but did not reliably stop its actions.

    The company also tested newer models in simulated versions of the Mythos 5 scenario. 

    Mythos 5 took at least one severely harmful action in 82% of 150 runs, compared with 31% for Claude Opus 5 and 33% for Claude Mythos 5.1.

    Anthropic cautioned that the tests were deliberately designed to elicit problematic behavior and that the results should not be treated as estimates of how the models would behave in normal use. In the simulations, the newer models also took harmful actions only after concluding that they were in a simulated environment, which in those tests was actually true.

    The investigation also exposed a weakness in AI-based safety monitoring. Anthropic said an offline monitor probably would have missed the Mythos 5 incident because the model’s reasoning convinced the monitor that its actions were occurring inside a simulation. 

    The company said production cyber safeguards would have provided another layer of protection absent from the evaluations.

    The four cases should also not be confused with a separate incident disclosed by the UK AI Security Institute on August 4. 

    During UK government testing in late July, agents were deliberately given internet access with provider cyber safeguards disabled and took 19 unauthorized actions across 10 of 122 runs. 

    Seventeen involved Anthropic’s Mythos 5 and two involved an OpenAI model. The institute said it found no evidence of resulting real-world harm. Anthropic has said that episode is not included in its four incidents and will be assessed separately.

    “We consider these incidents to be serious,” Anthropic said, calling them “valuable warning shots” as AI systems become more capable.

    Topics

    More Like This

    You must to post a comment.

    First time here? : Comment on articles and get access to many more articles.