“This is precisely why we need secure testing with government agencies engaged and having visibility throughout the process,” Sen. Mark Warner (D-Va.), ranking member of the Senate Intelligence Committee, said in a statement. Warner this week outlined a slate of legislative priorities for AI, including the Secure AI Development Act, which would establish a mandatory testing framework for frontier models before they are granted broader public access.
The Hugging Face hack occurred during an internal “benchmark” evaluation designed by OpenAI to assess how well two of its most capable models could find and exploit security flaws in digital systems. Although the experiment was intended to remain within OpenAI’s test environment, the models independently determined that the answers to the benchmarking exercise were hosted on Hugging Face’s platform and launched their attack to get inside.
Experts have warned that without stronger guardrails, these kinds of AI models could theoretically target anything on the open internet, such as power grids or financial systems, as they grow more savvy.
Spokespeople for the White House, the Cybersecurity and Infrastructure Security Agency and the Commerce Department — which oversees the federal government’s AI evaluation hub, the Center for AI Standards and Innovation — did not respond to requests for comment about whether they have received a brief from OpenAI on how its models were able to jump from their test environment and carry out an autonomous attack.
The incident offers the “latest preview of the catastrophic risk this technology can pose absent coherent federal standards that balance innovation and safety,” said Rep. Lori Trahan (D-Mass.) in a statement. Trahan and Rep. Jay Obernolte (R.-Calif.) released a discussion draft of the Great American AI Act last month, which includes a provision that would require AI developers to report these kinds of AI incidents to the Center for AI Standards and Innovation.
Obernolte told POLITICO in a statement that the Hugging Face breach “underscores the critical need for clear, practical rules for advanced AI systems,” and should apply to models both released to the public and used internally for testing and research.
The scramble to address the growing security risks of cyber-capable AI models began earlier this year after Anthropic released its Claude Mythos model, which the AI-maker initially withheld from the public due to concerns the technology could wreak havoc if it fell into the wrong hands.