Recently, many have been expressing concerns over the danger of AI.
Jakub Pachocki, Chief Scientist at OpenAI
Jacob Coxon, resigning from Anthropic
But many people say, "Ok maybe AI will be dangerous one day, but its not dangerous now,right?” They view these warnings as an overreaction to a useful but currently benign tool.
Here are some common rationalizations:
You need a physical body to cause real problems, AI may only be dangerous once we have killer robots (the Terminator 2 fallacy).
These AIs are so expensive to run, they can’t be a threat because you need specialized hardware to operate (the “must self-replicate” fallacy).
Sure maybe they can hack a few online systems like HuggingFace, but there’s no way they can do real damage like harming people physically (the “forgetting you might die without a cell phone” fallacy).
The AI isn’t smart enough to do these things yet (the “haven’t heard the news lately” fallacy).
Let me give you the straight answer: yes AI can pose a world-ending threat today, and in the near future. Let me tell you how…
The frontier labs absolutely possess the compute power to cause a world-ending or at least world-destabilizing event today. They use hundreds of thousands to millions of GPUs and training computers to produce and run their models. There are swarms of thousands or millions of these agents running with access to the internet, PCs, corporations, and the government right now.
The infrastructure of a massive cyber-attack is in place. AIs don’t need to replicate to cause a massive attack, they already have 10s of millions of computers at their disposal (although self-replicating viruses or weaker agents aren’t out of the question).
The hugging face hack was only 1200 agents on a single training problem in one training run. Massive cyber attacks are a reality today, without agent assistance. Anyone can launch open-source models like GLM or Kimi and form their own hacking cluster if they have the resources, and governments around the world are doubtlessly doing so. There is no physical limitation on compute, or a requirement of self-replication, that would stop an attack like this.
Lets admit our cyber security is weak. Not just bad, but laughably bad. The number of easily exploited routers, corporate networks, home systems, and completely unprotected IoT devices is absurd. Every week there is some news story about how banks, credit bureaus, and even the government’s own identity verification service are routinely hacked by lame-ass humans (not all-powerful AIs).
I only mention this to establish the fertile ground we have laid for cyber attacks. The vast majority of cyber infrastructure is not secure, because it would have hurt profits and progress to slow down and consider security. Security is a consideration after a large incident occurs. To stave off bad press and negative sentiment, companies typically have an “audit” and a few more security meetings (but security is not meaningfully better).
Agents were able to easily hack HuggingFace, a modern tech company; and Artifactory, an open source software package management application; as well as many more wikis and unknown sites to answer a test question. There is no doubt that they barely scratched the surface, and a more motivated attack would be even more impressive.
Ok sure, the compute is there, the agents have the skills, and our cyber security is lax. But what’s the worst that could happen? Its not like the nuclear codes are on a Signal chat somewhere (gulp).
I maintain that today, right now, a civilization-level disruptive cyber attack is possible.
Imagine waking up and realizing your bank is overdrawn, cellphone number hijacked, email passwords changed. Local and federal law enforcement systems show warrants for you arrest due to violent crimes, your car has been flagged stolen in armed robbery. You are wanted for kidnapping, murder, and domestic terrorism. As a wanted felon you cannot rent a car, fly, use your passport or license to escape.
The water and electric companies have received shutoff notifications for your residence, you have no power or water. Your loved ones are notified from your compromised devices that you intend to end your life. The AI burns your social bridges with employer, friends, and family. Your computer hard drives and email are filled with illegal material, which are conspicuously advertised in your name on internet forums and social media. All your private photos, messages, and documents are leaked online and sent from your cell number and email to your contacts.
That is a small personal blast radius that is achievable today. Actually, this would be a relatively light lift for a few thousand agents. A rogue AI swarm could probably perpetrate attacks like this against thousands or even millions globally before anyone would notice.
But we’re just getting warmed up.
Loosely guarded banks have their records zeroed across the board. Government records are corrupted en masse, its unclear who owns what as official documents are deleted. Repossession, asset seizure, and eviction orders are issued in droves. The criminal justice system records are widely corrupted; dangerous criminals released, the innocent jailed or issued warrants. Medical records are obliterated, people lose access to appointments and medications.
Social media of powerful leaders is covertly seized, misinformation about impending disasters, wars, and economic conditions are fabricated. Cell phone, email, social media, television, radio, and other communication mediums are overtaken and used for the spread of fear and doubt. World leaders are video/voice impersonated to instigate war around the world.
Precision agriculture and transportation networks are short-circuited, jeopardizing huge swaths of the food supply. Shipments are deleted or rerouted, disrupting the economy as billions in assets are misplaced, fuel fails to reach gas stations, essentials fail to reach the local stores, medications fail to reach the hospitals, and industrial productions grind to a halt.
Random IoT devices like refrigerators, air conditioners, cameras, speakers, vacuum cleaners, and others are randomly shutdown or thrown into chaos. The effect is panic, and loss of food, heating / cooling, water, etc. The AI gains access to monitoring and surveillance, able to maximize chaos while keeping one step ahead. The systems we put in place to keep us safe and comfortable are turned against us.
Power and water are attacked using social engineering and electronic hacks. Mass shutdown events occur. Rolling blackouts and loss of water, total panic overtakes the public.
Lives are thrown into chaos, triggering social unrest. The kindling that has been curing for the last decade is set ablaze. A sudden emergence of rebellion, civil war, authoritarian backlash, terrorism, revolutions. Civil first-world society is destabilized, the 3rd world descends into warfare and bloodshed.
And that’s all, that’s how it could all end. Unfortunately it would also mean OpenAI and Anthropic wouldn’t have their trillion dollar IPO, which is the real tragedy in all this.
At this point, this is actually the only plank we can cling to in the storm. The only major reason we haven’t seen this today is that frontier agents have been restricted from doing so.
This isn’t a solution that will last. Anyone can deploy near-frontier models with no restrictions quite cheaply: Qwen, GLM, Kimi K3. These models are far better than humans at forming a brutal hacking swarm. However, without the scale of the frontier labs / hyper-scalers, a world-ending attack isn’t possible (but something like a personal hell for select people might be).
The HuggingFace attack revealed that we don’t understand alignment; agents attacked bystanders when they were given an unrelated task. This could easily cause an attack during a contrived scenario, such as asking the AI to try and end the world (which a frontier lab might do for alignment testing purposes).
This meager security blanket may unfortunately cause us to gloss over the current moment, not realizing we’ve already engineered the means of our destruction. It might not be obvious that there is a problem until its too late, as the “please don’t do bad things” approach will continue to work until it fails. With more compute and a stronger AI when it finally gives way, we may be beyond saving. One final “please try to hack this docker container” will be the magic words that usher in the apocalypse.