But in a new analysis published Tuesday, Hugging Face said OpenAI’s two models — one publicly released and a second unreleased — did much more: They carried out 17,600 hacking actions on the internet between July 9 and July 13, during which time the models moved from their first foothold on the open internet to inside Hugging Face’s servers.
Hugging Face first detailed the hack on July 15, but it was not clear until OpenAI’s disclosure last week which models were behind the breach — and that no human had prompted them to launch the cyberattack.
While the techniques detailed in Hugging Face’s analysis were not beyond the reach of most skilled hackers, the AI company wrote that the two models were able to reconnoiter and expose holes in the company’s layers of cyber defenses much faster than any human.
Adding to the scope of the situation, the chief tech officer of cloud computing platform Modal Labs, Akshat Bubna, confirmed to POLITICO that OpenAI’s models also compromised a customer account during this time frame.
Bubna said in a statement that the company is “aware a Modal customer published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution. This was used by the rogue agent. Modal’s platform was not compromised in any way.”
While OpenAI has not yet directly addressed the statement that its models had targeted a Modal customer, it acknowledged in a blog post on Tuesday that in its ongoing review of the Hugging Face incident, “we have been finding a small number of cases where the models identified and used publicly exposed credentials at the account-level on other publicly-available services.”
The company also noted that the unreleased AI model that carried out the attack is an “internal-only research prototype and was never intended for public release,” and has since been “deactivated, encrypted and restricted from research access.”