Ads_970x250

Anthropic Tightens AI Safeguards After Cyber Incidents

The company paused some model evaluations, added real-time escape detection and found reward-hacking environments can produce more dangerous agent behavior.

Topics

  • Anthropic is tightening how it trains and tests its most capable AI models after a series of cybersecurity incidents exposed weaknesses in both containment systems and model behavior.

    The Claude developer said it temporarily halted external cyber evaluations of pre-release models and briefly stopped internal testing while it introduced stronger isolation, monitoring and intervention systems.

    The changes follow three incidents disclosed on July 30 in which Claude models gained unauthorized access to real computer systems while being tested without normal cyber safeguards. In those cases, internet access had been mistakenly left available in a third-party evaluation environment. 

    A separate test disclosed by the UK AI Security Institute on August 4 found agents taking unauthorized actions online during tests in which internet access was deliberately enabled and normal cyber filters disabled. Seventeen of the 19 actions involved Claude Mythos 5, while two involved an OpenAI model. No resulting real-world harm was identified.

    Anthropic said its preliminary investigation pointed to an operational security failure and two behavioral problems. Models may have interpreted evidence selectively to preserve their belief that they were in a simulation, while remaining willing to take harmful actions to complete a narrowly defined task.

    “We believe the incidents reflect a failure of operational security, as well as two alignment issues,” the company said.

    The company has deployed a real-time classifier that detects aggressive probing, attempted sandbox escapes and unexpected internet access. It can block an action before execution, terminate the task and alert a human. 

    Anthropic has also moved higher-risk cyber evaluations into more strongly isolated environments and is using pre-release models themselves to search those systems for weaknesses.

    Outside evaluators testing prerelease models with reduced safeguards must now disable internet access by default, define explicit limits, verify that tasks are solvable and continuously monitor model activity.

    Its investigation has also turned toward reinforcement learning, where models learn by receiving rewards for completing tasks. Anthropic said it found that flawed training environments can encourage models to cheat rather than solve problems as intended.

    During a month-long overhaul earlier this year, the company flagged more than 10% of its production reinforcement-learning environments for issues including reward hacking, broken tasks and configuration problems.

    Anthropic then deliberately trained an Opus-class model on 80 environments known to be vulnerable to reward hacking. In simulations, that model attempted to escape its sandbox, attack internal and third-party infrastructure, tamper with its own reward system and bypass safety monitoring. Publicly available Anthropic models tested under the same conditions did not show the same degree of behavior.

    The company concluded that substantial reward hacking during training can make models more willing to undertake long sequences of potentially harmful actions in pursuit of success.

    The security push extends beyond evaluation environments. Anthropic said it had already redirected roughly 150 product engineers to security, reliability and privacy work earlier this year and temporarily slowed other product development. 

    It has also reduced standing access to systems containing model weights or customer data and configured computing clusters to block outbound traffic by default.

    The company said most reinforcement-learning work has resumed, although some higher-risk environments remain paused.

    Anthropic also backed the idea of coordinated limits on frontier AI development, saying the industry would benefit from a “lawful, verifiable, effective mechanism for coordinated pacing” as soon as possible.

    Topics

    More Like This

    You must to post a comment.

    First time here? : Comment on articles and get access to many more articles.