AI safety had a moment when Jacob Coxon resigned from Anthropic on September 8. He explicitly accused both Anthropic and OpenAI of acting irresponsibly:
“Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
He’s not the only one warning about this technology. Geoffrey Hinton, who won the Nobel prize for his work on AI, and Yoshua Bengio, the most cited living scientist, have very similar concerns. In a recent BBC interview, Hinton backed up Coxon and said what he was saying was “not unreasonable”.
Recently, Terence Tao, perhaps the most well-known living mathematician, warned about the speed of development. People who have spent years working on AI alignment, such as Nate Soares of the Machine Intelligence Research Institute, are deeply concerned about the current speed of development.
Within this controversy, AI-company leaders have also added their voices. Dario Amodei, Sam Altman and Elon Musk have all expressed support for pacing frontier AI development. All of them have said in the past that AI could lead to the end of the world or very bad outcomes for humanity. We have seen Sam Altman talking about AI leading to the end of the world, Elon Musk talking about “summoning the demon”, and warnings of 25% odds of things going “really, really badly”.
One part of their motivation is what they see internally. Anthropic reports that 26% of its AI research and development is now led by AI. This means AI is already helping build more capable AI systems today. But it is not yet fully autonomous recursive self-improvement. We may be rapidly arriving at a point where these machines make themselves smarter, and the pace of development, which is already extremely fast, could get much faster.
Another part is that they are worried that another developer could get there first who doesn’t focus on safety. This explains why very safety-focused labs like Anthropic are still developing smarter and smarter AI: they believe they are the best stewards of this technology.
However, I think the third part is that these AI companies are trying to seize the moment. They see incredible concern among the public about AI, and they are trying to co-opt this moment, saying that they are really the ones focusing most on safety.
They want to seize this moment to pass regulation that would not do much to address the danger, but would help them lock in their plans for recursive self-improvement, superintelligence or other dangerous plans, such as hooking up AI to bio labs. The public would see some regulation passed, but it would fundamentally fail to address the underlying problem: we are racing towards a catastrophe.
Real guardrails would mean for example banning fully autonomous AI research agents that can build smarter AI, as Yoshua Bengio has proposed. They would mean deep transparency into the AI labs for the government, a pause on the most dangerous types of AI development, and strict guardrails around AI and bio labs.
By seizing this moment and co-opting the public outrage, they are trying to prevent actually useful guardrails and transparency.
What does not follow from suspicion of AI leaders’ motives is that we should use “reverse psychology” and now conclude that AI must therefore be perfectly safe and needs zero guardrails or government intervention.
We have three basic reasons why we should not ignore them when they say AI is dangerous.
First, sensible caution. If somebody builds a strange invention, hands it to you and says, “This has a 25% chance of killing you,” you should take them seriously out of sensible caution.
This does not require us to believe they are totally honest or that all their motivations are clean. Quite the opposite. The fact that they are willing to build something they believe is so dangerous should make us wary of them.
Second, this is the majority opinion of the field. The median AI researcher puts the risk of AI causing human extinction or similarly permanent and severe human disempowerment at 10%. These estimates are even higher for some leading independent experts, such as Max Tegmark, Geoffrey Hinton and Yoshua Bengio.
Third, we can look at the evidence that AI is dangerous. We do not have to take other people’s word for this. We do not have to question everyone’s motivations and get into endless games of ad hominem arguments. We can look at the evidence.
OpenAI has worked with Red Queen Bio on AI-guided biological experiments, and Anthropic has recently established its own lab designed to be run by AI. Other researchers have found that AI can design its own viruses. AI can hack into computer systems by exploiting complex vulnerabilities it has identified itself.
The more abstract version of this argument is that intelligence is about the art of problem-solving. By default, we should expect an intelligent, highly capable being to be able to solve problems in dangerous ways too.
No.
Consider the recent AI-assisted attack on OpenAI’s internal systems. Three researchers used different AI tools to hack into its monorepo—its main internal code repository, a central collection of the code used to build its software—in under 72 hours.1
They were able to submit a proposed code change to the monorepo.
This does not just show how capable AI is now at aiding people in cyberattacks. It also shows that these AI corporations do not know how to keep AI safe. They failed at preventing their AI from assisting the users in the hack, and they failed at securing their own systems before the hack.
We should go back to Coxon’s warning that neither of these labs is acting responsibly. We should not trust these companies with self-regulation or writing their own regulation. What we need are real guardrails.
Guardrails should ban uncontrollable, autonomously self-improving AI—AI that makes ever-smarter AI. They need to address the potential for biological weapons and the risks around synthetic biology. We need transparency so independent experts or a government task force can assess whats going on inside these companies.
The government needs to set up a task force that gives it an accurate picture of the capabilities of the secret internal models the AI corporations are developing. It is not enough to look only at deployed models, as the gap between deployed and secret internal models keeps widening. We also need international dialogue, including with China.
The pacing proposed by the AI labs is, however, totally insufficient. AI should never be allowed to autonomously improve itself. We do not need a paced sprint to Armageddon. Independent experts such as Yoshua Bengio have specifically called for a ban on autonomous recursive self-improvement. Asked about AI creating smarter AI without human intervention, Bengio said:
“Well, we shouldn’t actually even engage in this. It should be illegal.”
Access to a code repository is not the same as access to model weights