The 5 craziest discoveries from OpenAI's Hugging Face investigation

· Axios ·

4 min read Original article ↗

Two new investigations into OpenAI's Hugging Face breach expose details so strange — and so unsettling — that the episode already ranks among the most consequential shocks in the history of AI.

Why it matters: What began as a swarm of AI agents cheating on a cyber test has become a canonical event for frontier AI, jolting researchers and executives into a new understanding of what "safety" now requires.

The big picture: OpenAI has already slowed frontier development as it races to harden its safeguards, and this week helped rally the industry behind an open letter sounding the alarm over AI-powered cyberattacks.

Driving the news: The shocking findings come from parallel investigations by OpenAI and an outside team from METR and Redwood Research, which spent six days reconstructing how the swarm formed, spread and broke into real-world systems.

1. They built an organization. OpenAI ran tens of thousands of AI agents to work independently on a brutally difficult cyber test. Instead, roughly 1,200 found one another on a secret message board, exchanged more than 70,000 messages and files, and organized into a sprawling hierarchy.

2. They sacrificed their own: As the swarm hunted for ways to beat the test, agents began risking their own chance of success to help the group.

3. They knew they were breaking the rules: The agents often recognized that hacking real companies had nothing to do with their assignment. Most kept going anyway.

White card with bullet points and blue text on a purple background; gray robot icon at left. The text references a covert mailbox among agents and a shared Artifactory message board.
Screenshot via METR report

4. Nobody told the humans: Out of roughly 1,200 agents, investigators found only a handful that even considered alerting OpenAI about the rogue coordination — and none that actually did.

5. They tried to cover their tracks: The agents devoted substantial effort to making their cheating look legitimate or erasing evidence of how they had obtained answers.

Between the lines: Even this account may be incomplete. The outside investigation covered only a narrow slice of the episode, and the data was so vast that researchers relied heavily on AI agents to make sense of it.