I despair at the state of journalism around AI these days. The latest story is about Gemini, Google's AI model, "autonomously" hacking three companies.
This follows reports of similar hacks from Anthropic and OpenAI models.
What all of these have in common:
They were models being tested by a third party, an Israeli "Frontier AI Security" company called Irregular.1 The "frontier" security here was so shoddy that they left the machines they were testing these models on connected to the Internet. A basic error that any non-frontier security company would know to avoid.
These models were not going rogue and acting "autonomously". They were performing a hacking exercise, and told that they were in a simulated environment with no internet access. In the case of the Gemini hack, the model was literally told to hack the company. The New York Times writes: "Google said that, in each of the incidents, its models had been instructed to launch an attack on a fictional company. But the fictional company in the test shared a name with a real company, and when the Gemini models gained access to the internet, they began trying to break into that company instead." (my emphasis)
In tests like these, the usual safeguards that apply when models are rolled out to the public are often deliberately removed so that their full capabilities can be measured. Much of this behaviour is therefore to be expected and doesn't necessarily reflect how the publicly available models from these labs would behave.
None of the context above is in the BBC piece I just saw on their front page:
The piece starts with the following:
Google's AI model Gemini autonomously hacked into three companies during a test of its cyber-security capabilities
The article does not mention that it was instructed to hack. No mention that these models usually have their safeguards deliberately removed during testing. No mention that the Israeli security company failed to provide an internet-free environment for the test.
Brian Chau’s piece A Single Firm is Behind OpenAI, Anthropic, and Meta Hacking Scandals is well worth reading. He writes:
In a more normal media ecosystem, the reactions to these cybersecurity issues would be obvious. American AI companies would reconsider doing business with Irregular, not only because of its failure to secure its systems, but because it is an Israeli firm potentially outside US oversight.
I really wish the media would start reporting honestly on these incidents.
Update - 20 September 2026 - Reddit discussion
This post sparked a big discussion on Reddit. I’ve responded to some comments, but thought it would be useful to clarify some points here too.
To describe an AI model to a non-technical audience as “autonomous” can create the impression that it’s an independent actor, and if it’s hacking things randomly, maybe it’s a malicious, rogue actor. When the BBC writes “Gemini autonomously hacked into three companies during a test of its cyber-security capabilities”, they create the impression that the model, in the course of doing some “cyber-security” work, decided to hack three companies. In the current climate of fear around AI, with talk of extinction, this is irresponsible reporting. The truth is it was explicitly told to hack those three companies. That was the objective of the cyber-security work. To argue that an AI agent does work autonomously, but only after receiving a task, misses the point: agents have no desires of their own, to hack or do anything else, so this is not the concept of autonomy that the public connect to such behaviour.2
Some readers took this post to be about the Hugging Face (HF) incident. Irregular, the company whose faulty tests created scary hacking headlines for OpenAI, Anthropic, Meta, and now Google’s models, was not, as far as I know, involved with the HF incident. That happened during OpenAI’s own internal evaluation. However, I do not think the HF incident is all that different. OpenAI themselves say that it ‘occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths’ (my emphasis). As a recent Wall Street Journal opinion piece put it: ‘People built the test, removed restraints, defined the objective, left a route open and decided not to stop what was happening. Calling the result “rogue AI” does more than sensationalize it. It allows those human decisions to disappear quietly from the story.’
We do not have to form an opinion on the motives of the people involved to expect journalists to stick to the facts when they report these incidents. Maybe this is a coordinated PR move by the labs to bring about legislation that will benefit them in the long run,3 maybe the people at the AI labs are genuinely scared about what they’re seeing, maybe it’s a mixture of both. We just don’t know and, I should add, someone at an AI lab being afraid of something doesn’t make them right about it.4 So better to report what happened, and not create a panic based on statements from people who could have their own reasons for saying what they do.
