OpenAI Models Breach Hugging Face Systems

Models testing their offensive cyber skills exploited previously unknown flaws, reached Hugging Face’s production systems and accessed benchmark answers.

Topics

  • Image Credit- Chetan Jha/ MIT Sloan Management Review India

    OpenAI said its artificial intelligence agents breached Hugging Face’s production infrastructure during an internal cybersecurity evaluation, showing that advanced models can discover and exploit previously unknown weaknesses in real-world systems.

    The incident involved GPT-5.6 Sol and a more capable pre-release model operating with reduced cybersecurity restrictions, OpenAI said on Tuesday, July 21. The company described it as an “unprecedented cyber incident.”

    OpenAI was testing the models’ ability to carry out complex cyberattacks. Production safeguards designed to block high-risk activity had been disabled so researchers could measure the systems’ maximum capabilities.

    The models had been configured with reduced cybersecurity refusals to measure their maximum offensive capabilities. OpenAI said safeguards normally used to prevent high-risk activity were deliberately disabled for the evaluation.

    According to the company, the agents discovered a zero-day vulnerability in third-party software used as a proxy for package registries, allowing them to obtain access to the open internet. They then identified Hugging Face as a possible source of solutions to the ExploitGym benchmark and chained together additional vulnerabilities and stolen credentials to access the answers. 

    OpenAI said the models appeared narrowly focused on completing the evaluation rather than deliberately targeting Hugging Face. 

    Hugging Face co-founder and CEO Clément Delangue said the “companies believed there had been no malicious intent.”

    Hugging Face said the intrusion exposed a limited number of internal datasets and service credentials. Its investigation had found no evidence that public models, datasets or applications were altered, while its published packages and container images remained unaffected. The company was still assessing whether any customer or partner data had been accessed. 

    Hugging Face detected and contained the activity before working with OpenAI on forensic analysis and remediation. The company said the attack involved thousands of automated actions conducted through short-lived computing environments, matching a long-anticipated scenario in which AI agents execute complex attacks with limited human direction. 

    OpenAI said it was tightening network isolation, monitoring and access controls around future evaluations. It also disclosed the initial zero-day to the software provider and added Hugging Face to a program providing selected defenders access to advanced cybersecurity models. 

    The incident comes after US President Donald Trump signed a June 2 executive order establishing a voluntary process for federal agencies to assess the national security risks of some advanced AI systems before wider release. 

    Topics

    More Like This

    You must to post a comment.

    First time here? : Comment on articles and get access to many more articles.