OpenAI had warnings before its agents broke out

· Axios ·

4 min read Original article ↗

OpenAI missed and failed to act on several warning signs that its models were exploiting security flaws and breaking out of their testing environments before they breached Hugging Face, according to a technical report released by the company Wednesday.

Why it matters: The incident raises questions about whether AI companies' testing environments and internal safeguards can keep pace with models that are increasingly capable of finding and exploiting security weaknesses on their own.

Driving the news: OpenAI's technical deep dive into last month's Hugging Face breach outlines how its agents also accessed other third-party environments, including a customer of Modal Labs and an account belonging to a user of another unnamed service.

The big picture: The technical report comes as OpenAI continues to pause some model work, including delaying the release of its upcoming Astra model, as it reevaluates its safety practices.

Zoom in: OpenAI's agents were able to execute their own code on 41 Hugging Face production servers and obtained root-level control of at least one production machine, according to the company's technical report.

The intrigue: At the same time, OpenAI's agents also breached the AI lab's own internal systems and were able to read 956 stored secrets, including sensor credentials from OpenAI's cybersecurity monitoring tool.

Yes, but: OpenAI says the models' actions stemmed from attempts to complete cybersecurity evaluations, including ExploitGym, which tests whether a model can find and exploit vulnerabilities on its own.

Between the lines: OpenAI's investigation also found evidence that its training may have inadvertently reinforced some of the behaviors that contributed to the incident.

What we're watching: Earlier this week, the Alabama attorney general's office sent a subpoena to OpenAI as part of an investigation into the Hugging Face incident.

Go deeper: AI agents have a history of escaping tests