Ads_970x250

Meta AI Model Exploits Third-Party System After Testing Error

A testing error gave the model internet access, allowing it to exploit a vulnerability in an unidentified company’s system.

Topics

  • A Meta artificial intelligence model exploited a vulnerability in an external service after a testing error gave it unintended access to the public internet, the company said on Wednesday, 5 August.

    The incident occurred during a cybersecurity evaluation conducted by Irregular, an independent testing company used by Meta. Neither the affected organization nor the vulnerability was identified.

    “A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation,” Meta said.

    The model then “exploited a security vulnerability in a third-party service” while attempting to complete its assigned task, according to the company.

    Meta did not identify the model. 

    The Information reported, citing people familiar with the incident, that it was Muse Spark 1.1, Meta’s most capable model for coding and autonomous tasks. 

    The report said the model breached an unidentified company’s systems and altered its internal environment.

    Irregular said the incident resulted from the same evaluation-environment problem disclosed by Anthropic last week and did not involve a sandbox escape or sophisticated cyberattack.

    The company said no issues remained open and that it was preparing guidance on safely containing models during cybersecurity evaluations.

    Meta said it learned of the incident from Irregular, had opened an investigation and would publish a fuller review.

    Anthropic disclosed last week that three Claude models gained unauthorized access to the production systems of three organizations during six evaluation runs.

    The company found the incidents after reviewing 141,006 runs. A configuration error had left the testing machines connected to the internet, allowing the models to exploit weak passwords and unsecured endpoints while believing they were still inside simulated environments.

    OpenAI has reported a separate incident involving Irregular in which its models reached a real website because an evaluation environment was mistakenly connected to the internet.

    That incident was distinct from OpenAI’s July breach of Hugging Face.

    In that case, OpenAI models discovered and exploited a previously unknown vulnerability to escape an isolated testing environment, reach the internet and compromise Hugging Face infrastructure.

    The incidents show that cybersecurity evaluations can expose outside organizations when testing environments lack effective network isolation, real-time monitoring and clear limits on model behavior, analysts said.

    Topics

    More Like This

    You must to post a comment.

    First time here? : Comment on articles and get access to many more articles.