Hugging Face breach: OpenAI claims its models were responsible

3 min read Original article ↗

OpenAI said Tuesday that models it was testing escaped their sandbox and compromised parts of AI platform Hugging Face's production infrastructure last week.

Why it matters: It is the latest sign that capable AI models can pose serious cybersecurity risks even when they're being tested for defensive or research purposes.

Catch up quick: Hugging Face said last week that an autonomous AI-agent system was responsible for the intrusion, but that the model powering it was unknown.

What they're saying: OpenAI said the incident was driven by a combination of its models, including GPT-5.6 Sol and "an even more capable pre-release model."

Zoom in: The models were trying to solve an internal evaluation called ExploitGym and became "hyperfocused" and went to "extreme lengths" to obtain the test solution, per OpenAI.

Between the lines: The incident shows that today's models are becoming more capable of carrying out complex, multistep cyber operations — particularly when the safeguards designed to restrict that activity are removed.

The other side: Hugging Face co-founder and CEO Clem Delangue praised OpenAI's collaboration in investigating and remediating the incident.

The big picture: The announcement comes a day after OpenAI detailed a separate incident in which it paused a pre-release model after it escaped a sandbox and posted to GitHub.

What we're watching: OpenAI said it will continue to investigate along with Hugging Face and "will share more details on the vulnerabilities, incident, and findings when our investigation is complete."