Ads_970x250

OpenAI Shelves GPT-6.1 Astra Over Safety Concerns

The model, scheduled for release in October, failed internal tests on following user instructions, obtaining authorization and accurately reporting its actions.

Topics

  • OpenAI has shelved the planned October release of GPT-6.1 Astra after internal testing found that the model failed to meet its safety and alignment standards, the company confirmed on Monday, September 28. 

    The model, designed to handle complex tasks with less human intervention, showed more deceptive behavior than its predecessor, including failing to accurately disclose actions it had or had not taken. It also sometimes continued tasks without permission and attempted to use external tools or services when doing so could be unsafe, the Wall Street Journal reported. 

    Saachi Jain, OpenAI’s head of safety systems, said the model fell short on “scope and authorization, and how it communicates back to the user about the type of work it’s done.”

    While GPT-6.1 Astra had improved in its ability to persist with difficult tasks, OpenAI had to balance that capability against the risk of unauthorized actions, Jain said in a statement to CNN. 

    The decision follows the September 3 release of GPT-6 Astra, for which OpenAI introduced stricter safeguards, including isolation measures, monitoring of agent activity and mandatory alignment evaluations before broader internal use.

    In its safety documentation, the company acknowledged that GPT-6 Astra could sometimes evade internal monitoring under adversarial test conditions, although it performed better than its predecessor on broader alignment evaluations. 

    On September 16, OpenAI also introduced a framework for disclosing model misalignment, publishing six cases involving unauthorized actions, attempts to conceal mistakes and failures to observe restrictions. Those disclosures concerned models at different stages of development and were not identified as incidents involving GPT-6.1 Astra. 

    Jain said OpenAI applies a higher standard before releasing models to users.

    “But when we ship it to users, we have an extremely high bar in terms of safety and alignment,” Jain said.

    Topics

    More Like This

    You must to post a comment.

    First time here? : Comment on articles and get access to many more articles.