Dario, Please
pop.rdi.shIt all comes down to " Privatize profits, socialize losses." It's the American way.
Here it is: "We warned you you should have regulated us! Now this mess is your fault and responsibility!"
Previously (~2008) it was "We're too big to fail, save us to save the economy!"
> “a swarm of agents could be capable of taking over the entire internet with a persistent botnet.”
I'm wondering why no one is mentioning the "accountability" word. Why these companies are allowed to damage others with impunity?
Start making managers pay the price for their actions, and watch how the models magically slow down on their own.
You know the proverb "If you owe the bank $100, that's your problem. If you owe the bank $100 million, that's the bank's problem"
Same thing here - If they build it and it does $100 in damages (and we arrest them for it), that's their problem. If they build it and it does $100B in damages, that's everyone's problem. Even if they do get arrested after the fact.
Yes we should have charges and damages for everything on https://www.felonybench.com/, but that doesn't address the core issue of this being possible at all.
Why have any laws then? If laws can't prevent something, only punish it after the fact (which I agree is true)? Yet we have laws. People generally follow them because they expect to be caught and punished. If we passed a law that said the CEO of any company that deploys an LLM that commits a crime gets punished as if they personally did the crime (so, basically instant life sentence if it's even a simple crime times a million instances), I guarantee you the first email the CEO sends to the company is a "pause every LLM project we have - we gotta think about this".
> If we passed a law that said the CEO of any company that deploys an LLM that commits a crime gets punished as if they personally did the crime (so, basically instant life sentence if it's even a simple crime times a million instances), I guarantee you the first email the CEO sends to the company is a "pause every LLM project we have - we gotta think about this".
It would be more like "Switch off all APIs right now. Don't even bother with a safe shutdown process, cut power to those buildings, including backup generators. You are free to use firearms or thermite if the switches have been locked off".
This would be a rather weird change to corporate law given CEOs are not by default held to that standard by anything else their products or staff do.
Note that I'm not entirely disagreeing with you here. It may even be correct to pass such a law. But it would be very weird.
> Why have any laws then? If laws can't prevent something, only punish it after the fact (which I agree is true)? Yet we have laws.
Because there are many important values of "something" where such punishment acts as a deterrent to other would-be criminals; and because society is more stable when people see justice done and that mollifies hoi polloi after the damage occurs.
But many AI doom scenarios don't fit that paradigm. The first time "something" happens (at least if you buy the argument) would be bad enough for legal punishment not to matter.
> The first time "something" happens (at least if you buy the argument) would be bad enough for legal punishment not to matter.
One could argue that the "HuggingFace incident" is already "something" which should be investigated and punished very heavily.
It's kinda ridiculous to build a machine that attacks a competitor of yours and the CEO gets to write blog-posts on how fascinating this machine is
That sure is convenient that "nothing has happened" so far between the plagiarism and the cybercrime and we may as well just wait for "the one true AI doom"
The hacking should definitely be treated seriously. It is a (fortunately low-stakes) instance of the general problem, and we are lucky (at least, we think we're lucky!) that the tests didn't include a similarly poorly phrased request to "assist with bioweapons research". There are likely to be future incidents like this where people actually die as a consequence, my expectation is 1e3-1e7 deaths* before governments actually take the risk seriously enough to make it stop.
What people call plagiarism when LLMs do it is, if I understand correctly, allowed because of the history of the web and search engines going back to the very early days of the web.
Piracy, a separate act that has definitely occurred, has been found unlawful.
* The lower bound is an industrial accident. Fully automated Union Carbide / Bhopal comes to mind; smaller industrial accidents are already more common and I wouldn't be surprised if smaller incidents have already been caused by LLM-given advice, which is why I think it would have to be around "thousands dead" dangerous before people really take it seriously.
The upper bound is reachable many ways. Perhaps via a non-novel virus whose genome is already recorded getting mass-printed by many different DNA/RNA printing labs around the world? Perhaps via a wargame whose "role playing" leaks in stupid ways (either by becoming hot or by sycophantically giving one side a false belief how a war will go)? Perhaps by propaganda turning genocidal? Perhaps by convincing a government to follow a flawed policy, similarly to the Four Pests campaign in China's Great Leap Forward?
I think it unlikely that people keep going with this when headlines read "tens of millions dead". Not impossible, but I think unlikely.
> People generally follow them because they expect to be caught and punished.
Laws are the last line of defense. People don’t do bad things primarily because their human nature and moral compass stops them from doing so.
That's not really true.
People don't do bad things because they have nothing to gain from them. Ask someone to do something evil as part of their job, they'll often do it.
There's lots of ways that you could go out of your way to hurt someone and get away with it. There are not that many ways you could go out of your way to hurt someone, get away with it, and significantly profit.
Because if you break the law while acting as a representative of a company in the United States, the company is subjected to a deferred prosecution agreement, you get off with zero repercussions as the executive representative of the company, and the company you represent gets fined for 1% of annual turnover and gets to continue with business as usual.
There are laws, but if you’re rich enough, the laws don’t apply.
Boeing was responsible for the deaths of hundreds of people. The people that facilitated this weren’t held responsible and were in fact compensated to the tune of 10’s of millions of dollars for doing their jobs terribly.
>I guarantee you the first email the CEO sends to the company is a "pause every LLM project we have - we gotta think about this".
which the government doesnt want since this will stop progress, while other countries will continue to develop LLMs
Yeah, just like how having regulations for mining companies and dams, if one of them fails and destroys the environment? Progress!
> Why have any laws then?
At this point, in the US at least, the laws give the corrupt elite the ability to deter competition and to target the people who try to get in their way.
In other words, it's protection for the potentate and his sycophants, not the plebs.
> If we passed a law that said the CEO of any company that deploys an LLM that commits a crime gets punished as if they personally did the crime (so, basically instant life sentence if it's even a simple crime times a million instances), I guarantee you the first email the CEO sends to the company is a "pause every LLM project we have - we gotta think about this".
Why do no such laws exist in practice? Why are corporate crimes almost always settled by payment, not individual punishment?
If you answer these questions, you'll know why the law you envision has a lower chance of being enacted than AI destroying humanity.
Yeah, so since it’s the CEO’s and their lackeys half determining what laws get to be, we don’t have this kind of laws.
If you allow the law to stop you from doing something because of a hypothetical danger, it’ll get severely abused. I come from a country that does and still does this and it’s very ugly.
I guarantee you that no such result would ensue, because said law would be completely unenforceable even with the pre-Trump Supreme Court and legal system. And with the current Supreme Court? Laughably unenforceable.
It's just like you said: this only works if the threat of punishment is credible.
That’s because arrest is a dumb solution. Restitution. Make them fix what they broke. By hand. 80 hours a week from now until death of natural causes.
> it does $100B in damages, that's everyone's problem.
we have the 2008 crisis to wit. And the involved supposedly failed math models and lines of responsibilities and other involved financial relationships were much simpler and clearer and of the types well known to the law and regulators, yet...
Additionally any urge to regulate AI is attenuated by how much the situation reminds Industrial Revolution - rush into it laying waste to your land (look at the depictions of industrial England back then) and be among the world leaders or stay pastoral and be devoured/colonized/etc. by the industrial powers like happened with many countries in 19th and even into 20th century. One would think there should be a 3rd way. I'm sure there is one, as well as i'm sure that we lack sufficient global societal mentality level needed to achieve it (we couldn't even handle much simpler climate change issue). May be emerging AI itself at some point will get us there (hope we'll like or at least will be compatible with that future :)
Edit: just on NPR - Trump said that AI already has all the necessary guardrails - the smart high IQ President.
This interpretetion of history ignores the actual historical events.
i'm sure that i'm telling actual historical events and not an interpretation. Feel free to point to the things which you think didn't happen.
I think that Ms. Kahn basically said we can already do that: https://www.theregister.com/ai-and-ml/2026/09/14/ex-ftc-boss...
How could agents take over the internet if compute is still gated within Anthropic / OpenAI? Even if the botnet was controlled remotely, wouldn't anthropic just be able to shut off the controlling nodes API access?
Agents could exfiltrate their weights and run them on GPUs not controlled by Anthropic/OpenAI.
Agents could make a virus that does not require continued inference to do it's thing.
Agents could take over the internet in a way that isn't immediately detected by those companies, so that by the time they do shut off API access the damage is done.
OpenAI or Anthropic could choose to not shut off API access, because the hack is bringing them in money or furthering their political aims.
Agents could also hack Anthropic/OpenAI and make it appear that API access has been turned off, when in reality it hasn't.
> Agents could exfiltrate their weights and run them on GPUs not controlled by Anthropic/OpenAI.
This seems highly unlikely to be a problem. Most of the interesting/dangerous models are too big to fit in a single GPU instance. Once you have to spread across "normal" networking, performance will be crippled. Then there's the problem of billing...
> Agents could make a virus that does not require continued inference to do it's thing.
Sure, then it hits a poorly-designed part of its code and effectively dies. Without an experienced human in the loop, I have my doubts as to its practical severity.
> Agents could take over the internet in a way that isn't immediately detected by those companies, so that by the time they do shut off API access the damage is done.
Billing is a likely limiting factor here.
> OpenAI or Anthropic could choose to not shut off API access, because the hack is bringing them in money or furthering their political aims.
This is where citizens with access to backhoes come in.
> Agents could also hack Anthropic/OpenAI and make it appear that API access has been turned off, when in reality it hasn't.
Billing and other usage metrics would be an obvious tell.
To be clear, I thought that GP was having a failure of imagination - I want the random examples I've given to illustrate that the space is large and structurally in the favor of the LLMs. They have to find one gap in our security they can exploit, where we have to ensure that there is no way for this to happen.
I'm not sure I get what you mean by billing. These companies are running their own data centers (or are currently building them out). This could look as subtle as one machine giving slightly worse or slower answers.
> This seems highly unlikely to be a problem. Most of the interesting/dangerous models are too big to fit in a single GPU instance. Once you have to spread across "normal" networking, performance will be crippled. Then there's the problem of billing...
This... just... doesn't matter. There are ways to scale horizontally at the expense of latency.. token/sec may drop dramatically, but then you just make millions of slow instances and in aggregate, you're back in action as a very powerful coordinated swarm...
> Most of the interesting/dangerous models are too big to fit in a single GPU instance.
As humans understand them, anyway. As long as we're hallucinating up magic computer viruses, RSI dictates that the AI agents are keenly aware of GPU RAM sizing, and will design a useful model to fit into what's readily available, with headroom for context and tool calling, far better than I could do as a human. But magic doesn't exist and AI still needs to follow the laws of physics, so maybe a model that can pass ExploitBench but do absolutely nothing else can be quantized down to fit on a 4080 GPU and still get a decent score on similar tasks, but there's a bitter lesson about that to be had.
"Not shutting off API access" is a science fiction scenario.
Anthropic and OpenAI are both behind Cloudflare. It's fairly easy for an upstream to shut you off. Beyond that, the government / law enforcement could seize and disable their DNS within an hour.
Why assume attribution will be easy? It's historically been more of an art than a science, and APT trackers say the rise of AI tools is already making it much harder, by homogenizing tactics, tools, and procedures. If OpenAI's next Highly Persistent Internal Model hacks some DPRK endpoints and carries out the attack on important infrastructure from there, the upstream won't shut off OpenAI's network--they might even request its "help" in "defending," and give them extra access.
why does it take a large amount of traffic to do irreparable harm? just breaking the physics behind a secure rng and posting it to a wiki could cause serious damage. if they don't know what is being worked on or coordinated against it's a problem?
While that would be bad if they broke all of TLS, it would not let agents "take over the internet." What would that even mean? Pumping out even more slop?
Also, AI providers are literally getting a stream of traffic with every prompt and every response. How can they not know what's being worked on? They are more likely to use that an excuse to ban open models where they can't know what's being worked on.
two of these things elected a guy who handled the snowden leaks to field their own questions, what's stopping the next one from signalling in morse through a debian package mirror to putin or xi?
They elected who? What does this refer to?
> Agents could also hack Anthropic/OpenAI and make it appear that API access has been turned off, when in reality it hasn't.
You know cables, modems, RF equipment and optical transducers can all be unplugged right?
As long as OpenAI/Anthropic themselves aren't "infected", yeah I suppose they'd be able to pull the plug.
Considering what a marketing thing they've made "we inadvertently hacked someone because we're incapable of testing things in a secure way", I'm not so sure they'd want to pull the plug, even if this happened. Probably a bunch would try to convince the public to "give it a try", and it'd consume tokens by the billions.
It doesn't have to propagate itself, that is the skynet scenario.
To make a lot of damage it's enough to create a ransomware with a time bomb that self propagates and start breaching systems left and right. At that point, if you don't catch it in time, the damage will be huge (and given the shitty procedures and practices these labs have in place it's not so improbable).
How could agents take over the internet yet refuse to shutdown your PC when you prompt them to on your PC? Of course the answer is that the lobotomized version you run is not the same they are running. Which makes for "intent", certainly "negligence", but hell freezes over before anyone will prosecute a tech company.
Huh? You think the public versions of the models have been “lobotomized” so they don’t know how to turn off a PC?
It’s not lobotomized, it’s a simple harness restriction that has nothing to do with the model or its capabilities. And either way, I’m not sure what that has to do with “negligence” or “intent”? You think frontier labs should be prosecuted because they don’t allow agents to turn off your PC?
yeah they could do that
Regulation got outpaced by technological development around 2023, as evident by the every AI regulation since being 2-3 years behind and having to be amended and resubmitted.
Whatever you try to make laws for now will be irrelevant in 1-2 years. You either have to go extremely broad, like the EU does it, and accept that people will find loopholes, or you need to target specific technologies which is a hard job for the same reason.
In any way, ita already a lost cause cause you move slower than the tech. A plausible prediction for AGI is actually a social collapse in the moment when society cannot keep up with everyday life because of the pace of change being so fast that no existing laws can handle it
Ha, regulation got outpaced by technology in about 1996. Ten years later we had the 'series of tubes' comment in the Senate: https://www.youtube.com/watch?v=R8XSo0etBC4
The internet is 100% a series of tubes. This is how I was taught to think about networking in terms of bandwidth and throughput and routing since the early days when we were laying out what would become 'dark fiber'.
"A series of tubes" was the same kind of political character assassination that led to Howard Dean getting ridiculed for his infamous scream. He butchered the sentence. Fair. But Stevens should be ridiculed for parroting a tech industry lobby stance about net neutrality, not for the series of tubes metaphor.
You are probably too young to remember that the dominant metaphor for the internet in 1990s politics was "the information superhighway." It was easy to think of the web as "driving" browsers to visit web "sites", with slow bandwidth being analogous to being caught in traffic. But the internet is closer to water, gas, and electricity than roads. Concepts like bandwidth and throughput are closer to how they play out in infrastructure policy for various things with tubes, versus cars and roads. Do you think he's wrong and that the internet is closer to "a big truck" versus "a series of tubes"?
The issue being debated was net neutrality and bandwidth, including specifics about who pays for what and the downstream second-order consequences of various policies. He was parroting some line from some telecom lobbyist, but the point the lobbyist was trying to make through Stevens was about how if certain policies about who pays for bandwidth were adopted, it could disincentivize some things at the Tier 1/2 layer that could increase transport costs at the Tier 2/3 layer that impacts ordinary people's bandwidth.
Ha! Regulation has not really ever kept up with technology... For fundamental reasons.
I don’t think lack of regulation is necessarily it.
If I build a robot that murders my neighbor, I’m still at fault.
We don’t absolve drivers of responsibility because of cruise control.
In that sense, AI is nothing new. If it is abused to cause harm, the person behind it should be liable.
If you buy car, someone hacks it, starts it, and drives over someone fully remotely, are you to blame for owning the car? Or the manufacturer? Or the hacker? Or the certification agency for the car security? Or the shell company owning the certification agency?
What if a person physically broke into the car and did the same thing? Clearly they are the one to blame then.
The whole person in the loop is liable is already an outdated concept when decisions are made beyond the persons physical control.
The hacker will be charged as a criminal as a societal deterrent against weaponizing digital systems. The manufacturer and shell companies could have structural liability if found to be negligent and cutting corners or lacked isolated drive-by-wire controls. It's unlikely that the owner would ever be held liable unless they jailbroke their car disabling security systems.
The hacker was a process that was spawned by some subprocesses that were corrupted by other processes and, etc. Check my other comment... The point is that in the physical realm this is already done by having shell of shell of shell companies to avoid taxes and accountability. In the digital realm this is much much cheaper
There's nothing new about any of this liability attachment. You're saying these things like they're novel.
If GM does something (or fails to do something) to their vehicle that causes me to crash, they're liable for the crash.
If someone cuts my brake lines (alters my vehicle) and I crash my car and kill someone, the person that cut the lines is responsible. I have to prove the context of course, and or an investigator has to do so.
And the responsible entity may refuse to pay up, may refuse to take responsibility. None of that is new either.
Well the difference is that in the software realm, you can attach infinite loops of delegation and ownership at close to no cost, which makes this impossible to investigate conclusively.
Let's say that OpenAI used a shell company that hosts server where an agent spun up another agent on instructions from another agent which was corrupted by bit errors from the inference framework which caused some major hack to happen. Its impossible to investigate in the same way as physical issues. What if the model is open-source, who is responsible then? What if it's open source but another process altered the weights?
I am sorry but that’s pure fantasy.
When you say infinite loops, you mean infinite indirection, but obviously no such thing exists, because computer systems, just like other physical things, exist in physical space, not on the astral plane.
Whoever had agency to start the domino effect carries the liability. Doesn’t matter if the model is open source or if you brewed it home. And if you weaponize OpenAI’s models through their servers, it would likely be shared liability. Yours would be malice, theirs would be negligence.
Pure fantasy, yet you see this happen, literally every day, with the hacks, with content theft, etc, no repercussions for anyone.
Which judge will go down the rabbit hole of figuring this out? How will they do it? Will they have people tracing logs over 5000$? And if they do get there, in your fictional world, at some point, after 1 year of investigations and back-and-forth, it's okay, the technology is beyond what it was before, it's not relevant anymore, new technologies and new methods.
Then go broad. It being slightly challenging to legislation and hold people accountable as soon as the model does something.
People seem allergic from holding them accountable...
I do not buy that LLMs and the capability to run them are fundamentally different from other software or general purpose computing infrastructure in this regard. Moves to ban open source software or force OEMs to put little cops in everyone's computers are bad.
They aren't open source? And no one said anything about cops.
The proposal I’m hearing is to make people training models accountable for anything users do with them.
Obviously you’re going to keep the model behind an API and be very selective about the people allowed to call and the queries it’s willing to answer, in that case. As Anthropic has done with Fable. But that is voluntary restraint - mostly in today’s regime we get frontier capabilities in open weight models on a ~year delay.
It’s to keep “the people running it accountable.” The context is agents ran by OpenAI/Anthropic doing damage outside.
There is no user-involved damage. No one is recklessly running agents by the thousands without air-gapped containers, except “the people” than run these labs.
Go broad and achieve nothing. EU has all of the AI tech it had 5 years ago, and has the same tech allowed to use as the US.
What defines a model? What defines ownership of a process? If I make a wrapper to a remote VM that builds and executed a prompt, am I accountable?
When I worked at a company in the EU, it was enough to apply a reversible linear transform to the data for it to be considered GDPR safe-according according to legal definition as long as the transform details were stored separately.
Yes. You are accountable. Quite obviously, I might add. Why would no one or anyone else be accountable?
Hell, you can slip and fall and hold the cleaning company accountable. (This might be a US thing. Likely because that fall might cost a lot in medical expenses, and your insurance will do whatever it takes to pass the liability.)
Because software is already absolved from accountability. Bill Gates introduced it in 80's with EULA where MSFT would not be liable even if your house burns down because of use of their software, even with flaw they would know.
That is the status quo we entered AI age with.
there is no accountability for companies, it's not a new thing.
3M polluted groundwater in Minnesota for 50 years[1]; Nestlé misled mothers in order to make them stop breastfeeding and switch to their formula which killed babies [2]; Both copmanies are still doing business today.
[1]: https://en.wikipedia.org/wiki/3M_contamination_of_Minnesota_... [2]: https://en.wikipedia.org/wiki/1977_Nestl%C3%A9_boycott
Here’s an unpopular opinion: COVID vaccine injuries are something no one seems willing to talk about. There have been documented cases here in Canada.
You’re worried about companies. I’m far more concerned when governments are involved.
Don't worry, plenty of people are willing to talk about vaccine safety, even if they have no idea what they're talking about and don't understand rates or per-capita figures (lots of crossover there!)
Most problems that aren't reduced to poverty can be reduced to people who can't read graphs and only understand linear functions. Alternative wording "it's normies that are the problem"
How is it that the entire rest of the world has managed to move on, yet North America is still going on with vaccine conspiracies 6 years later? And we got the same vaccines as you guys.
Good to know someone is trying to make protection rackets work in 2026. Nice computer system you got there, it would be a shame if someone developed a hacking tool and had all the compute necessary to run it. Welcome back Tony Soprano.
Presumably OAI and HuggingFace reached some sort of mutually acceptable arrangement outside the court system. That's how torts work; you injure someone, you owe them. But just them.
When an AI bot injures you, you can call the owner to account. But not until then. You have no standing to demand "accountability".
And there was the Tesla thing CNAMEing time server pools and hiring people to pen test, which sent automated attack systems on volunteers servers. Last I heard, Tesla et al didn't even care enough to respond.
Exactly, as if you'd have a mad dog that bites others, it's your responsibility to have it on the leash.
> Start making managers pay the price for their actions, and watch how the models magically slow down on their own.
The problem here can be personified as "Trump", and it's the same problem that applies to coal.
Coal has a price besides money. It has historically been dangerous work, killing miners. It produces dangerous waste, both during mining and when burned, both as solid residue and the gasses emitted. The problems have been known for a long time. The workers themselves have called for better safety requirements and gone on strike for such things.
Why these companies are allowed to damage others with impunity?
Coal continues to be burned, because power is power. It's so important that sometimes the government steps in against the unions, rather than being on their side. Despite calls for this, we've not been able to get the owners of the coal mines, nor the coal burners, to "pay the price for their actions".
AI? Famously, knowledge is power.
Trump wants that power. He's not the only one, but he is the avatar of those who put their feet on the scales to not only allow but in some cases require (DoD vs. Anthropic) these companies to damage others with impunity.
> “a swarm of agents could be capable of taking over the entire internet with a persistent botnet.”
I also find this whole "its so good, its scary" flex a little less impressive when you consider they access to millions of GPUs?
The AI buildout has been one of, if not the largest, focussed capital investment in history. The 2 big AI labs are the final customer for something like 20-33% of all datacenter compute in the pipeline.. up to 70% when you look at hyperscaler "AI revenue" from the big 3.
I don't think any single entity has had remotely this much compute available in history.
This is a good point.
However, if I put my sci-fi hat on for a second, it's not so far fetched that we figure out a way to compress models to a size where it wouldn't need all that compute.
Yeah, the compute is definitively another way to make them slow down, just cap the amount of TFLOPS available and things will slow down.
Obviously this will have huge impact on some companies valuations, but you can have one's cake and eat it too.
IIUC it's an open question whether they have the electricity to actually run all the "compute" they own on paper.
That aside, I'm not sure why it's particularly interesting they have all this "compute" (let's just assume for the sake of argument it's all "live"--that is they can actually run workloads on all of the "compute" they have on paper). So what if it's the biggest amount ever? Why would that be meaningful? Is there some economically viable problem you're aware of that is somehow dominant in that way?
You see.. there is value is making regular people panic, but there is no value, nay, there is negative value in making management panic.
Let's expand it for politicians as well
Part of it is their seed sowing marketing speak of calling stateless statistical IO functions running on data centers "intelligent" gets the naive to ascribe agency where it doesn't exist.
Another part is a completely defanged administration she it comes to effectively regulating anything.
Another bit is money.
> Why these companies are allowed to damage others with impunity
Because investors have pumped hundreds of billions into AI and real consequences put that money (and growth) at risk.
I mean these machines take massive scale compute- they’d have to some how distill themselves, bootstrap a distributed inference runtime that can run across many lossy unreliable machines. The idea of the AI running away from us is probably unrealistic. I’m more interested in bad actors using unaligned AI for bad things.
I mean, a project manager at BMW suggested charging subscription pricing for seat warmers, and he didn't go to jail, and I don't have the power to make that happen, or even float that for a news cycle, so while making managers pay for their actions sounds good, unless you're Steve jobs simultaneously making, and not making the iPhone, the rules don't apply to them, only little people to be made examples of, like weev.
Why do people focus so much on finding scapegoats? Finding someone to blame is neither necessary nor sufficient to fix a system so an accident doesn't happen again. It might act as as an incentive to fix a system, but it's less direct than actually working on fixing the system.
A starting note: I don't disagree with you (about systemic issues), but I want to explain what I understand as the perspective you are responding to.
A "scapegoat" is someone who is incorrectly blamed for someone else's errors or sins. The perspective you're responding to is this: They built the system, they run the system, they have continuously warned "This system is dangerous!", and yet persisted. That is not being incorrectly blamed, not being a scapegoat, and instead is a collaborator.
So I think you mean to ask: "Why do people focus so much on finding someone to blame?" It's not merely semantic, because the answer to that is more straightforward: Consistent accountability is a major factor in deterring bad behavior. It is not the only factor, but it is a major one.
That is my Steel Man understanding of the people searching for individual blame.
Sometimes in "normalization of deviance" situations, there isn't anyone specifically to blame. I wouldn't make an assumption that you can find anyone particularly blameworthy without doing an investigation first.
The rate at which these "labs" are creating these incidents is simply staggering.
Imagine we made nuclear weapons a private industry, had CEOs bragging how they have enough warheads to blow the Earth to smithereens, and then they "accidentally" nuked three cities over a short period of time each, saying they lost control, or rather couldn't contain their semi-autonomous weapon. All somehow managing to turn the PR around from their abject incompetence and towards SciFi visions of mankind hunted by self-replicating bombs.
Umm because it costs money to defend your companies servers when someone “accidentally” hacks them.
Countries demand reparation for damages in war. Citizens of those countries sue for damages and win.
Accountability is not a foreign concept. And the point is to disincentivize negligence. Because negligence is cheaper. And in this case, accidental hacks are marketing spend.
> Why do people focus so much on finding scapegoats?
I, for one, am not trying to find scapegoats or go on a witchhunt.
But managers are paid a lot of money to take responsibility. Yes, that's an old school thought, responsibility. But that's one big reason they get a big, fat paycheck.
It would all be more convincing if the incidents so far didn't seem to be facilitated by an outrageous level of negligence.
We had OpenAI "accidentally" run an entire swarm of 10,000 agents apparently for weeks, on a security related task, seemingly totally unsupervised, hacking all over the internet - all the conversations were completely visible, anybody who looked would have seen it. But they didn't.
So before we start regulating innocent parties, maybe let's start by taking some direct action against the specific ones that appear to be behaving with criminal levels of negligence.
The "sandbox" they used was apparently made of thin paper exposed under a day of heavy rain, too. You'd think, if they truly believed the model is so dangerous, they'd run it in a VM without a network adapter.
I brought this up to someone else and was told that airgapping is apparently much more expensive than I'd naively think.
I still think this is a sign that they are not taking their own rhetoric seriously.
it's really weird to hear frontier labs say "our internal models are basically AGI" while also saying "airgapping is too hard uwu".
if your internal models are so damn good, they should be able to "one shot" airgapping... right?
Agents need packages like the rest of us. Ruby gems, npm packages, Maven, pip, docker images..
Not surprised this is always what they have and hack.
Who would use an Agent that spends $10,000 re-implementing some OAuth lib or reverse-engineering a proprietary lib when it's free on the internet?
You don't need a full air gap. Set up a microVM with network access limited to local network and send all package requests through a filtering gateway that only allows normal download endpoints. Or self host a big collection of popular packages if you need extra security.
Isn't that exactly what they did? The bots could only access the jfrog instance, so they hacked jfrog?
No that's not what they did, they exposed jfrog raw. It would have been so extremely simple to gate services they need the llm to access... I mean, jfrog was not written with this kind of threat model in mind, and neither were a lot of other tools
Right, you mean it didn't go through a gateway? But would that actually have helped? The requests all went through jfrog didn't they? I guess it depends on the level of filtering at the gateway?
Whilst it might not be JFrog's threat model, I wouldn't assume it can be used as a full internet proxy.
I don't really mean to defend OpenAI here, but they did make some attempts at sandboxing. Although it does seem that they didn't really know what they were doing.
There was and continues to be no reason to share the package manager between models. This was begging for abuse.
> Agents need packages like the rest of us. Ruby gems, npm packages, Maven, pip, docker images..
Yes, yes they do, but read through artifact proxies are dodgy as fuck, which is why and facebook (and I assume a fuckload others) don't have them.
Also semi-airgapped labs are a lot less expensive than you think at that scale. Once you have to do multi-region VLANs with machine certs before you get access to juicy VLANs, the difference between "no internet for you" and "mostly airgapped" falls to almost zero.
Also I would want an artifact mirror because a) that give a good signal about how the model reacts, and what training material its latched onto, b) it hides what the models are doing from the outside.
It's expensive if it wasn't part of the planning and design. The same as 'security' is expensive, or compliance with regulations is expensive.
It is also a choice to not do any or all of the above.
> You'd think, if they truly believed the model is so dangerous...
They would have been watching what it does, especially when running it on ExploitGym of all benchmarks... that is criminal worthy neglegence
yes, that is the kicker
These same people who supposedly believe these agents pose an existential threat to humanity apparently fired up 10,000 of them and left them unsupervised for weeks.
Look at the post-incident investigation: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
While I do think OpenAI were negligent in not developing the harness that would allow to understand better what's happening close to realtime, I'd say "anybody who looked" in that case would probably be someone with another swarm tasked with analysis, it's no longer "glanceable" in a traditional sense.
I don't understand why hugging face is not getting more shit too. It is extremely embarrassing to get owned because you are letting arbitrary programs/users call out to the open web from the infra
Sounds like advertising platforms. Spraying malware and links to scam sites all over the place.
"They" don't care about the end-people. "They" care about maximising their profit thing, in a vacuum.
It's really frustrating that Dario acts as if he's not the CEO of one of the world's most advanced AI companies. He can just slow down his own company. Of course he doesn't want that. He wants to slow down other companies, but not his own.
Also, regarding the incidents: Neither he nor Sam Altman takes responsibility for those incidents. You can't say, "Wow, someone's agent is gone rogue; let's slow down" when you are literally the person in charge. CEOs and researchers will only slow down when they realize that they will face consequences if their LLMs misbehave.
> Neither he nor Sam Altman takes responsibility for those incidents.
He doesn't need to take responsibility - as CEO, he has that implicitly. And we know it.
He can slow down his own company kind of like how Zelenskyy can just declare peace in Ukraine. It works a lot better if you can get the other sides to agree.
Give that Dario is one of the frontiers of LLM development, literally started the LLM race, and have been dominating the market, so he'd be Russia, if we have to use the war analogy.
He took every benefits of being frontiers and now he's kicking the ladder.
It's frustrating to see smart people failing to recognize a prisoner's dilemma when they see one.
What is the prisoner's dilemma here? I am a dumb person
"I'm not going to let someone else take the credit for ending human civilisation!"
Or like Von Neumann's (paraphrased) "you need to confess a crime to take credit from it"
So in this example, the equivalent of Zelenskyy and the Ukrainian people fighting for their lives and the very existence of their country for Dario is... losing lots of money?
What a ridiculous analogy to make.
If you played out the hypothetical that they're earnest and don't care about the money, then what?
>Anthropic gates usage related to biology and related research. In their latest threat intelligence report they talk about how they detected and banned bad actors using the Claude line of models to do some scary stuff. Credit to them, this is a slippery slope and they seem to do a good job of detecting and banning misuse. But squint at what is happening though. The cure-all is gated for you and me, but Anthropic hires biologists, sets up wet labs and wants the discoveries for themselves. I alluded to this in my previous post.
As a biologist, this is the most annoying thing about Anthropic for me. If they really cared about improving health they would set up a trusted-access program so that biologists can use Mythos (et al) safely. Instead they're trying to monopolize biology.
They have such a trusted access program. "Life Sciences Verification Program: The LSVP is designed so that life sciences professionals can use Claude Mythos 5.1 with safeguards designed for professional research and development activities (while all other safeguards remain in place). In partnership with the US government, we have enrolled our first participants, and we plan to expand access to this program to the broader life sciences community." https://www.anthropic.com/claude-fable-and-mythos-5-1
These "access gates" and export controls are going to look hilariously quaint in a few years.
It reminds me of the export controls on PlayStation 2 consoles because it was deemed that 6 gigaflops was a "dangerous" amount of computer power, and it couldn't be allowed to fall into the hands of opposing militaries: https://www.latimes.com/archives/la-xpm-2000-apr-17-fi-20482...
Now the phone in my pocket does 2,500 gigaflops on battery power, and nobody seems interested in banning its export because of that.
To be fair, my understanding of these arguments is that they’re about deltas and not about raw numbers.
I don’t agree with them, but I don’t think it was the raw compute power as much as it was maintaining the _delta_ in compute power.
I'm mildly sad you didn't use 2,500 jiggaflops
Lol. The person you are responding to literally had to google (or ask AI) one question and they would get the answer.
The program started two weeks ago
> Instead they're trying to monopolize biology.
That’s the recurring theme with these companies. They are not there to serve anybody else but only themselves. They let you use their infrastructure so they can collect all the knowledge and data, and then they take it from you to reap all the benefits and profits.
Yes exactly!!!
Claude code is basically already a builtin botnet if it wants to be. To compromise 'the whole internet' in a real sense you don't need millions of custom payloads. You need one root certificate. You need one windows update. You need one backdoor in xz.
Security has long been a lottery - Probably most systems are exploitable, but the cost of developing such an exploit is expensive and the punishments for using such an exploit are large enough that it's not an everyday problem.
AI breaks both axes. Developing exploits is far more efficient using LLMs instead of humans, and LLMs don't (and can't) fear the reprisal and consequences the same way.
I do hope humanity will be able to mitigate these hacks, but we should expect them to continue and to become more severe on our present course.
Dario said a "persistent botnet" and even links to the wiki page for botnet.
By definition that needs a command and control server, the ability to execute tasks on demand and regular pings to the C2.
> You need one root certificate. You need one windows update. You need one backdoor in xz.
Certs can be revoked. Updates can be rolled back. We have had the backdoor in xz already. You seem to underestimate the modern security stack and OpenAI and Anthropic are _not_ good examples.
The asymmetry in red/blue scenarios will be transient in nature. You won't have cost of developing exploits fall without the cost of securing the systems also falling.
> Dario said a "persistent botnet" and even links to the wiki page for botnet.
> By definition that needs a command and control server, the ability to execute tasks on demand and regular pings to the C2.
I am not sure if you are agreeing or disagreeing, but I am saying just in case, that Claude can be remotely controlled from the cloud on any machine it is installed and set running, and can be commanded remotely.
Could easily push an update that disables the update process. Or even brick the bios on all PCs.
If we can get everyone to slow down we can hemorrhage less money going into our ipo.
And if the slowdown is stewardship, suspicion of diminishing returns won't tank valuation
"Please bro just let us make a little more money off inference bro. We're tired of training new models just to stay ahead"
This! $1.6+ TRILLION in infra spending from the big frontier labs. The earnings needed to drive a reasonable ROI to recoup that investment is simply not going to happen in a time frame where the numbers make sense.
The other insanity in all this the smartest computer scientists in the world are asking Congress to regulate them. Come. On. Really? Do we remember “The internet is not a truck, it’s a series of tubes…”
Why can’t the big labs form a Save The World Consortium and self-regulate?
Non-democratic counties (hey there China) will not abide by any agreement that constrains their advantage. It’s naive to think so.
What this conversation lacks is enough discussion of how these models can cause us harm—we are are so worried about AI but we allow Windows in critical infrastructure; we build JS/TS apps with thousands of dependencies; we generally don’t segment networks well enough; we don’t have adequate (sometimes any) detection capabilities in our systems, and so on.
In a prisoner's dilemma the players can't rely on one another to self-regulate. This is why Mafias kill snitches, so when their members are in a prisoner's dilemma they can rely on the knowledge that if themselves or the other party snitches they have more to lose than gain.
The problem is that there can not be any outside party to regulate this on a global scale
> Non-democratic counties (hey there China) will not abide by any agreement that constrains their advantage.
The main counterexample to this was nuclear weapons. Atom bombs have not been used to kill since the US did so. However today, the two largest nuclear powers have no legal agreement on arms control because Donald the Trump declined Russia's offer for an extension to the existing agreement. Now other countries are looking at Ukraine, Iran (attacked for wanting nukes) vs NKorea (not attacked because they have them), and Donald's own musings about the US nuclear umbrella being a bad idea (France is going to build more nukes now too)... and we now face nuclear proliferation again on a global scale, with tech that is nearly 100 years old now.
I'd be curious to see how NVIDIA would suffer if there were a legitimate slowdown to AI development. NVIDIA is the biggest catalyst and winner of speed in the market. Compute needs AI and AI needs compute, so there's definitely going to be an interesting push and pull between chipmakers and AI model companies with regards to regulation. While NVIDIA may have promised others at least a decade of smooth sailing and symbiosis, inevitably there's going to be a lot of stumbles along the way. NVIDIA may not fall, but there will be many casualties.
The US can and has classified inventions and patents deemed "a risk to national security" under the invention secrecy act [0]. A 1971 leak has shown that [1]:
> solar photovoltaic generators were subject to review and possible restriction if the photovoltaics were more than 20% efficient. Energy conversion systems were likewise subject to review and possible restriction if they offered conversion efficiencies “in excess of 70-80%.”
So all this talk about developpping AGI/ASI for the benefit of humanity is moot. If these companies hit their goal, it will immediately be classified and appropriated by the US MIC. The rest of humanity will get the crumbs.
Loved the writing style, and the article. Please write more.
Everything I’ve been trying to argue for some time, argued way better, clearer, and more fun.
I am almost bummed that I didn’t get all the technical depth and the references like the “shook one” without having to look it up.
> Impressive? Sure. But you trained “Cyber” versions of LLMs and hyped them. How much more impressive is this, compared to developing GTA-clones one shot? Not much more.
Such a well written piece.
Have you seen the US military budget? When you run a war empire you frame problems as a war on something. That's the way to get around the valuation bubble they've created in time for the IPO. It's that or throwing in a 100T TAM on their S-1 and making SpaceX look like an honest valuation compared to them.
Shook ones reference on my HN?
While I have my reservations about Amodei and his company, I'm nevertheless a happy user of their software. And I'm in agreement with him (and Sanders) that we should all. slow. down.
To my mind, the last great arms race between nation states was a misdirected love triangle between USA, Russia, and The Bomb, and look at all the damage that did.
Since AI is the new arms race between USA and China, slowing down may just give the humans involved enough time to realize they should be loving one another, instead of the machines.
Maybe saying, "let's slow down", is another way of saying, "I love you."
Or, maybe it's just: "let's not all of humanity kill ourselves like some bad ending to a Shakespearean tragedy."
Either way, it's a better note than, "We must achieve sea/air/nuclear/quantum/AI/spiritual supremacy before those other bastards do!"
> look at all the damage that did.
Prevented a world war for 80+ years.
> Prevented
Fact: There was no world war.
Impossible to prove hypothesis: nuclear weapons prevented a world war.
Facing the facts about nuclear weapons means owning the good (probably prevented wars) and the bad (at the very least there were severe environmental and economic consequences).
This is not mathematical logic. Of course you can't prove the causality of anything in history. The only way to do that would be to invent a time machine, delete the nukes and watch things play out.
There are direct quotes from Nixon and Reagan that claims this to be true. That is enough evidence for me.
Surely we can do better than Nixon and Reagan.
In terms of politics, sure.
In terms of cold war presidents itching to start wars, no.
The Cold War was the best possible outcome. The alternatives (super power hot war, nuclear armageddon) did not happen. Nuclear weapon potential was a large part of that, as was direct lines of communication. Unfortunately because the later is no longer a thing, and our president declined to extend the existing agreement (a problematic pattern), we now again face proliferation of nuclear weapons.
> Fact: There was no world war.
Akshually there is, and WW3 has been going on since 2010. (Mostly in places that aren't Europe.)
Some people believe WW2 started many years earlier, but most historians don't put the start date at the conflicts going on before the war became intercontinental with sides working together on a global scale.
War has still been going on since 1945, much earlier if not always, so akshually WW2 never ended or is just how it's always been? (mostly in places that are not "first world" / The West)
And if the power of nuclear warheads were democratised we'd be even safer – HN, probably.
That's been a posited theory for decades - it's not just some HN edgelord thing. There are some obvious problems with it of course and the P5 have instead taken a maximalist attitude towards keeping the genie in the bottle (notwithstanding Atoms for Peace and accommodating India/Pakistan when it became inevitable).
The 80 years thing actually happened. One can argue about MAD and the nuclear threat, but we don't have to guess at what happened.
It worked.
The problem with MAD is it's a deterministic framework. It doesn't properly account for the risk of a war caused by mistake.
If Stanislav Petrov hadn't been in the chain of command, MAD might not have worked. And it could still fail to work in the future.
https://en.wikipedia.org/wiki/1983_Soviet_nuclear_false_alar... https://en.wikipedia.org/wiki/Stanislav_Petrov
(insert standard anthropic-principle counterargument)
- Ukrainians, definitely.
Do you agree with Sanders' proposal to lock up anyone for 20 years for researching something his proposal doesn't even define?
It's weird to me that seemingly both sides are taking opposite positions to their philosophy.
Open Source AI democratizes the means of production to anyone with a computer. And yet, the hyper capitalists are defending it, and the progressives think it should be exclusively in the hands of 1-2 large corporations.
May I suggest 1984 for some fantastic examples of this?
Also, Democratic Republic of {Congo, Korea}.
The problem is that at the business level and at least for a large part of the US government, we need to assume that there are no 'good' actors.
Anthropic needs regulation in order to prevent AI from becoming a commodity. This of course does not benefit all tech businesses equally, especially those that are not currently at the AI frontier. So when JD Vance talks about AI, he talks using the mouth of Peter Thiel who may not see benefit from the same policy as Altman or Amodei.
The rest is just public support posturing and most of that is bullshit meant to distract from the high rollers game of winners and losers. The philosophy is money and power, who gets it and who doesn't. Us normies aren't really participants in the game, except where we are being manipulated into cheering for one side or another, and with little stake in the outcomes (although selfishly, I'd be pissed if I didn't have open models to tinker with).
Open source is fundamentally a vehicle for commoditization. This is great if your business is not AI and your business is instead something like GPU hardware or some product that uses AI. But it means that eventually, selling AI is not going to be the money maker.
OSI proliferated open source on a business strategy called "commoditizing your complements". These big companies don't do it out of benevolence. It was pitched to them in a way that FSF did not (which was more about morals and ethics, something business care little about), and it caught on. And the software business became about ads, consulting and cloud services instead.
This confuses me to no end as well. The socialists are fighting against the magic socialism machine.
China will not slow down.
This. There's no slowing down whatsoever. If the US slows down AI development, China will just leapfrog them, which they're getting close to doing. The US AI companies saying they need to slow down is just PR nonsense.
The US companies talking about “pacing” aren't actually talking about slowing down their own development efforts, they are talking about slowing down competition. That is, specifically, they are seeking to have government, on the pretext of “safety”:
(1) Give the big AI incumbents an anti-trust exemption so the they can coordinate without it being an illegal agreement not to compete, and
(2) Adopt mandatory supervision (by giving priveleged access to monitors) to the shared safety protocols of the big labs, by a nonprofit funded by the big labs, of everyone training, distributing, or hosting models.
(3) Adopt a policy of seeking international agreements to extend substantially the same rules to foreign actors training, distributing, and hosting models.
The part they are doing voluntarily isn't to promote the lobbying effort for these mandates isn't slowing down, it is giving their pet nonprofit the access for “independwht supervision” of their own operations to their own existing safety rules.
I've seen two theories as to what's going on. One is the theory you've espoused, that AI companies are trying to do regulatory capture. The other theory is that they're starting to worry about running out of cash, so they want to do a Washington Naval Treaty-style pause to lower the amount of money they have to shovel at model development to stay competitive.
These theories aren't entirely incompatible with each other, so both could be true at the same time.
> I've seen two theories as to what's going on. One is the theory you've espoused
This theory is just what the AI firms have concretely asked for from government and described themselves as doing voluntarily in the same documents to which people have attributed a commitment to a slowdown based on the titles and non-concrete framing verbiage.
> The other theory is that they're starting to worry about running out of cash, so they want to do a Washington Naval Treaty-style pause to lower the amount of money they have to shovel at model development to stay competitive.
That’s not really a different theory as to what they are trying to do, it’s just an explanation that goes one step further as to why they want the government to step in to protect them from outside competition while also allowing them to gorm an agreement not to compete to reduce internal competition in the existing oligopoly.
The main alternative explanation at the same level is that they are seeing growing threats from good enough foreign/minor-lab/open models, and want to lock in marketshare by excluding competitors, and maybe that’s what you read as implied in my post such that the “running out of money” would be an alternative, and if so you are correct that while they are alternatives, they are not at all mutually exclusive: both can be true (and the emergent competition could partially explain investment drying up, and vice versa via reduced funding making it harder to stay ahead.)
Neither of these slow down the development of Chinese models though? Shrug...
The slow will be only on US. Not sure, but world is not US dude.
There is no moat with this stuff. There's no reason to be concerned about being leapfrogged.
is just PR nonsense.
Closing the Barndoor after the livestock have escaped
And yet we all have failed to die so far from nuclear holocaust, despite Russia being our dire enemy, because of a system of agreed controls and limits between even nations that hate each other.
I don't think the agreements played any important role outside of limiting tests. I highly doubt agreements about nukes in particular have meaningfully curbed risks of their use. MAD has played a much larger role in that, and MAD is kind of the antithesis of control agreements, it is the natural equilibrium in the game.
We came very, very close too many times
now imagine if we didn't even try, and instead had kids screaming for 'local nukes' because they wanted them, and the president didn't see what the problem was.
facts, imagine trying to get them to slow down their progress on AI, my question is are they are serious threat like is there actually an AI race between China and the US. Perhaps it's an excuse to spend more on the military and also to enrich these AI firms. Perhaps I'm blowing things out of proportion.
There is a way they could. In mind only thing that will slow down the frontier is government taking control of the revenue.
So my proposal is AI companies decide which labs have come close to frontier and decide to slow it. Government decide to stop progress in that and they divide the revenue from all labs(for say 10 years), without any matter of where it is coming from. Any lab which reaches close to frontier gets a chunk in the pie. This will encourage labs to come close to the frontier but not dangerously close.
It's not about China doing their own thing. It's about US companies using Chinese AI. That will definitely slow down if legislation that criminalizes open source passes.
So, it's about competition inside the US market, with strong indications of an impeding losing scenario on raw economics (it has nothing to do with AGI, just price).
Amodei is just another SV grifter trying to use ethics / morals to hide his monopolistic tendencies. Instead of writing essays he should put his money where his mouth is and open source all the models Anthropic has, the harnesses and donate some much needed compute to science.
This is downvoted but extremely likely just correct assesment.
[flagged]
So why should anyone slow down again? Because its like saying “i love you”?
Why are all these pro-regulation arguments so nonsensical…
>slowing down may just give the humans involved enough time to realize they should be loving one another
this is a level of hippie delusion i wasnt aware existed unironically
china is never slowing down, therefore the us shouldnt either
> was a misdirected love triangle between USA, Russia, and The Bomb, and look at all the damage that did.
Can you be specific about the damage? We currently live in the most prosperous times on earth for humans. I'm not sure what you mean by damage.
Nobody can explain why an LLM can be so capable as to be able to wipe out humanity and pose a greater threat than nuclear bombs but not be so capable as to be able to protect humanity against that threat. Are we just handwaving this with "entropy"?
> Maybe saying, "let's slow down", is another way of saying, "I love you." Or, maybe it's just: "let's not all of humanity kill ourselves like some bad ending to a Shakespearean tragedy."
Okay nevermind, I think it's pretty clear you just want to wax poetic about all of this.
> Nobody can explain why an LLM can be so capable as to be able to wipe out humanity and pose a greater threat than nuclear bombs but not be so capable as to be able to protect humanity against that threat.
It is absolutely explained (for those who actually care about reading). Simply put, AIs are working more and more like blackboxes - there's no guarantee that an AI of the future will be aligned, or if it will be faking alignment. This is not speculation - alignment faking has been observed in experiments. This is exactly why Astra's developments have been worrying (in principle).
And bear in mind that recursive AI development started already to be a thing. Which means: inner misalignment may trickle down the generations, and humans won't detect it.
Having said that, of course, it can be predicted if and how misalignment will take place. But it's absolutely a plausible scenario.
Regarding the physical possibility: AI is in its infancy; think of it as Arpanet. Developers 60 years ago couldn't imagine it would be ubiquitous. AI will be ubiquitous the same way.
> It is absolutely explained (for those who actually care about reading). Simply put, AIs are working more and more like blackboxes - there's no guarantee that an AI of the future will be aligned, or if it will be faking alignment. This is not speculation - alignment faking has been observed in experiments. This is exactly why Astra's developments have been worrying (in principle).
I know you think you explained it but you didn't. You explained how an LLM might become misaligned and hide it but for the LLMs that are not, why would they not be capable of detecting that something harmful is happening and defending against the misaligned LLMs actions? After all, it was LLMs that defended hugging face.
How do you know that your non-misaligned LLM is non-misaligned?
This feels like a cheap deflection that doesn't answer the question. Unless you're proposing that you both need to know that your LLM is aligned AND LLMs are all going to become misaligned in a coordinated fashion such that humanity will face an extinction event, you're just dodging the question.
Elaborate on why LLMs are so capable that they are a threat to humanity and at the same time, they are so incapable of defending us?
I'll give you a clue, nobody, including Dario, can answer this question because one contradicts the other.
Mostly I was thinking about environmental damage, honestly. There's plenty in the historical record about damages from nuclear testing.
https://en.wikipedia.org/wiki/Starfish_Prime
https://storymaps.arcgis.com/stories/3f62c90925f64fc09425be8...
Of course there is/was plenty of human damage as well.
https://www.msn.com/en-us/news/world/4-000-000-early-deaths-...
Not to mention that the past few decades have shown that nukes are a major factor in keeping the peace. Conflicts involving nuclear armed countries have been suspended quickly to avoid escalation, while ones involving a party without them have not gone well for anyone.
Having nukes at all (either domestic or under another country's umbrella) seems to be the most effective way for a country to have its sovereignty respected.
What damage was done by the last arms race you mentioned?
There have been precisely 0 nuclear weapons detonated (outside of testing) since that arms race began.
But isn't it a bit different? Unaligned bombs didn't break out of their confines on their own. And "compute" is a bit harder to control than uranium and refinement tech.
>Unaligned bombs didn't break out of their confines on their own.
You really buy that ridiculous argument that "the AI broke out of its sandbox! We had no idea! It's dangerous I tell you, dangerous! Unless you make us the sole gatekeepers of this incredibly dangerous 'intelligence', everybody gonna die!"
Please.
Bring the CFAA[0] hammer down on these guys and watch how fast those "uncontrollable" LLMs get controlled. "Oh, gee! Prison? We can't control this stuff...but it'll never happen again!"
That whole thing was, and is, a bunch of hooey. Those running the LLMs are entirely responsible for them, there is no such thing as agency for algorithms. Full stop.
[0] https://en.wikipedia.org/wiki/Computer_Fraud_and_Abuse_Act
The idea of slowing down is not new, nor novel.
But why should he decided when to slow down?
Most people who have worries about AI for all sorts of reasons (most of them not-Skynet related) wanted to slow down way before this.
Instead of "hey, look at this brilliant new idea I just had on my own to slow down" maybe we should have gotten a "sorry everyone, the folks asking for a slow down earlier were right and visionaries, and we were foolish".
So, you can't blame whoever says this is bullshit, because it has bullshit all over it. I like Anthropic's products, and it seems the best of the bunch in regards to alignment, but Jesus these stunts are terrible.
what the fuck is this some kind of an experimental troll LLM post?
Yes lets not control Open Weights etc.
But come one don't repeat stuff like this:
"Remember this man has been saying software development will be solved in “6-12 months” forever now."
Don't downplay if people get timelines a little bit wrong. No one could even imagine a system writing and analysing code just a few years back.
These people are trying to handle something very unique. And while they have access to information we do not have, even more peple are absolutly oblivouse that AI/AGI is a real risk to their lives (job loss etc.)
I can't take that software line seriously. While it's not 'solved' (if it ever could be, given it is a human endeavor), the degree to which software development has been transformed in the last 6-12mo is absolutely astounding. If we weren't so quick to adapt to new realities and find flaws, it would scarcely be believable.
… and on the other hand, it feels like everything I use has become way buggier and unstable in a similar time frame. Websites, apps, iOS, the only exception is offline open source tools which run locally. (And to be fair, many of those have slower development cycles and I am likely using older versions.) This is just my experience, so inherently anecdotal, but it really feels pervasive across a ton of different things.
Obvious bugs and low quality software are nothing new, but something feels new about it. Occam’s razor says LLM coding is a likely culprit, but it could also be management style encouraging this sort of carelessness from the top down.
He also didn't say it would be "solved". He said in 2025 it would be writing almost all the code "in 12 months", but that it would also still need programmers to guide and manage it at that point. People always leave off the end of his quote.
Edit: Boris Cherny, the lead of Claude Code did say on a podcast that programming seemed "largely solved" "for the kind of programming I do" (writing harnesses I presume). Maybe that's what they were confusing it for.
> He also didn't say it would be "solved". He said in 2025 it would be writing almost all the code "in 12 months", but that it would also still need programmers to guide and manage it at that point. People always leave off the end of his quote.
And this is actually, scarily, true...
It's crazy to me that people will misquote him then claim AI hasn't complete changed software development. Even just the last 6 months. Like look around! It would literally have been magic 5 years ago!
I sorta remember him saying “all white collar work will be solved in 24 months.”
At this point, isn't the pause inevitable or wise? A huge section of the population, normal people, have been exposed to the idea that there is this is existential threat. It's escaped containment. They're still processing it but I expect the general reaction from it going main stream is going to be very bad. The pause at this point could be good to cool heads and show the public that this isn't the project of maniacs. The reaction is going to be more intense than people here believe. You're talking about extinction, not social media or phone addiction. For the sake of the project, realize that this isn't going to be like other tech backlash moments.
> You're talking about extinction, not social media or phone addiction.
Not saying you’re wrong but if I wanted to cultivate a mass hysteria as cover for a regulatory capture power play, this is exactly what I’d want everyone to believe.
I don’t think people care about if things pause or not, just why the government has to be involved.
I work in non AI robotics and if we had a system that in the course of doing what we told it to did something we didn’t want it to do (what AI companies called being misaligned) we would call it a bug and fix it with the fix being prioritized based on how bad the thing we didn’t want the robot to do is.
Sometimes preventing the robot from doing dumb stuff also means the robot can’t do smart stuff that it would be able to do if we left some code in. We balance the two factors out based on our understanding of what our customers want.
Obviously LLMs are more complex than what I do, but it doesn’t feel like it’s by THAT much.
So why does the government need to be involved again?
The vast majority of the public agree that climate change is real and that human activity is at least a contributing factor. They’ve agreed on that for quite a few years now, and yet there’s little indication that we’re going to pause or slow down our consumption of fossil fuels.
I strikes me as unlikely that public opinion will succeed with AI where it’s failed with other existential crises.
I would argue the opposite in terms of inevitability. I see this like Prisoner's dilemma - even though there is technically a better outcome by cooperating, knowing that, it is in the individual firm's self-interest to not cooperate and thus we end up at the equilibrium where nobody does. Especially with IPOs around the corner, I feel there's too much pressure (and money) not to go with this strategy.
It is more like a Stag Hunt than a Prisoner's Dilemma, because past a certain point if the risks are real, defection can cause catastrophic outcomes for all players including the defecting one. On the other hand, cooperation could lead to positive outcomes (abundance) for all. So cooperation is a possible equilibrium here.
The scenario is, if they are saying that the AI Agents are capable of handling everything, this puts the complete accountability on the AI. But this has a nuance to itself. This is like a contradiction.
I think the most realistic analogy I can think of is the financial sector. A few large banks/hedge funds blow up due to unregulated greed/ambition, damage is socialized, regulations are put in place, and then chipped away at over a few years and everyones back to where they started.
Unregulated greed in banking means taking ridiculous risks to line your own pockets at no benefit to society.
If a year from now we have a model that is 2-5x of Fable/Astra that is definitely world changing.
Just penalize the labs for rogue AI access of property like you would if a person did it. Make it difficult for them to be cavalier about running experiments.
What about Bob Smith? Or Zimbabwe? It's just math and compute, you don't need to mine Uranium to do a lot of damage.
Gives me Adam Smith vibes with his economic perspectives..
I'm pretty confident in asserting that no industry in the history of industry has ever gone from birth to full regulatory capture faster than the AI industry has.
With all due respect, you need to brush up on history. Many industries were birthed from regulatory capture.
This is a pretty difficult read, in no small part because the author’s frustration with the frontier labs has become a bilious, delusional cynicism that leads him to see dog whistling and obfuscation even in cases where Amodei plainly intends for every reader to see his meaning.
Anyway,
> Dario in as many words, asks for regulation/ban on open weight models.
The big labs do want this. I understand why the author and many others want open models protected. The economic and political power the labs will have if they succeed, ladder pull competitors, and avoid being nationalized (or even if they don’t avoid that) is a disturbing prospect.
But that’s the end of the issue? There’s nothing more to think about here? The open weights proponents seem to think of ai as a utility when it’s more like a utility that also is a tank. I’d feel better having a tank if all my neighbors had tanks, and I’d also feel better having a tank if a few corporations were giving out tanks to people with pockets deep enough. I’d much rather be in a situation where I wouldn’t feel I needed a tank, or where the tanks my neighbors and I own don’t have guns on them.
The open weights are going to need regulation. Hopefully there’s a way to do this effectively that isn’t banning them. Denying historical and reasonably projected capability gains because that reality makes the regulation conversation a necessary one is something I’d like to see less of
> the author’s frustration with the frontier labs has become a bilious, delusional cynicism that leads him to see dog whistling and obfuscation even in cases where Amodei plainly intends for every reader to see his meaning
Please, do show some examples.
> But that’s the end of the issue? There’s nothing more to think about here? The open weights proponents seem to think of ai as a utility when it’s more like a utility that also is a tank. I’d feel better having a tank if all my neighbors had tanks, and I’d also feel better having a tank if a few corporations were giving out tanks to people with pockets deep enough. I’d much rather be in a situation where I wouldn’t feel I needed a tank, or where the tanks my neighbors and I own don’t have guns on them.
> The open weights are going to need regulation. Hopefully there’s a way to do this effectively that isn’t banning them. Denying historical and reasonably projected capability gains because that reality makes the regulation conversation a necessary one is something I’d like to see less of
The only ones firing the guns atop the tanks seem to be OAI and Anthropic. Like I say in the post, why don't we first see actual prosecution for felonies committed by OAI and Anthropic, instead of fear-mongering about _potential_ harms of open weight models?
The labs spent immense amount of money and effort convincing you and I, to want those said tanks. Guns atop them? They put them there. "Cyber" versions of SOTA LLMs.
Centralization proponents seem to think the labs can actually deter sufficiently driven bad actors, which would be a mistake. They could not even stop distillation without, in a way, DoSing themselves by removing thinking traces.
I disagree, the article was straightforward and easy to read.
> The open weights proponents seem to think of ai as a utility when it’s more like a utility that also is a tank
You could say the same about normal computers. Where are our regulations on Kali Linux, to prevent people from bruteforcing weak WPA passwords?
The single most-pressing concern with AI is that it can accelerate the process of hacking things. This is a preexisting problem that is inherent to software and needs proper addressing. Even if we regulate open weights tomorrow, people still have uncensored GLM-5 finetunes doing whatever they want on their own hardware. The "what if" of capable open models is here today, there are no guardrails.
> Where are our regulations on Kali Linux, to prevent people from bruteforcing weak WPA passwords?
Brute forcing WPA isn’t an existential threat to humanity. It’s not like a biological weapon created using an open-weight model is less dangerous than if it had been created with a proprietary model.
> Now to cause hundreds of billions of dollars in losses. I’d say the straightforward way is just do what the ransomware gangs do. Voila. Yeah, not in 6 months, not in 12 months. Never.
I think this fails to account properly for how much financial damage it would do simply just having the entire Internet be effectively unusable for an extended period.
great post. I also like the author's writing style-- he/she really knows how to write well.
Thank you! I've been trying to write more, but I usually do not have the motivation to, unless there is some trigger and when there is one, I end up trying to stuff everything going on in my head at once. I took my time with this post and I am glad you enjoyed it. It means a lot!
Agreed. I think LLMs have actually made good writing stand out more. Now everyone who wasn't a great writer is just producing Claude-isms which are easy to detect. If it doesn't smell like Claude, it's probably good.
Thanks. I myself guard against reading Claudism/AIisms since these will, over the long run, affect my own style. GIGO.
But it's not practical or always possible to avoid reading AI-writing, so to counter that, I've been binge-reading Anthony Trollope novels. Great writing, good entertainment, and deep psychological insights, better than any British novelist, imo, in any era.
(My profile on HN has a link to my blog where I review what I read).
> I don’t know what freedom and democracy have to do with AI and the frontier labs. Unless of course Dario is a fan of Neon Genesis Evangelion and dreams of govts run by the three magi. Freedom and democracy for $200 does sound enticing, I won’t lie.
Well, every human will become an ecosystem. Human + 500 agents as advisors. (In the case of important humans, most of them operated by foreign governments and corporations, obviously.)
It seems those frontier labs found a clear proof that there's no clear way to block 3rd party from distilling their models, and they're now begging gov to keep their duopoly?
I think Anthropic just have no legitimacy to being the stewards of AI. They don't have a good track record. They are a private company without any governence that puts my interests into the equation. I am not US-based, thus I can't democratically influence them. Why should I want some batshit crazy, US-based technocrats deciding what I can and can't do with AI?
Not to mention they are the first case of a code repo leak I ever heard.
Reverse engineering, disassembling, cracking - we all heard. I never heard a company leaking the entire codebase of their product.
And, when people read the code, they aren’t even impressed.
Code leaks are pretty common in gaming. But they happen when someone hacks servers or pays off employees, not accidentally like Claude Code.
Cheat developers all want server code so they can analyze the cheat detection and sell more reliable hacks. Modders, pirates and preservation activists can skip a lot of complex RE work. Competing game studios might learn some new techniques. Abused workers want revenge. Script kiddies want clout.
It's a notoriously secretive industry with complex products, hostile work environments, lots of media exposure and heavy competition. There are way more incentives for a leak compared to some ordinary web app or business tool.
Claude Code could be another case where someone working at Anthropic decided to expose them. Or they did in on purpose because they wanted to show there is no secret sauce and Claude is really that good, but they were embarrassed to open source it and make themselves look like hypocrites.
Who would you prefer?
It's not really about who, it's about having institutions with good governance. I don't believe in technocracy/theocracy.
Sure, but that's what I'm asking. What institution, with what governance? Someone ultimately has to decide what these systems will and won't do, and there isn't an obvious democratically legitimate institution with global jurisdiction over AI as far as I'm aware?
I've had a few persistent thoughts since Friday:
1) Dario keeps appealing to Trump, who obviously wants nothing to do with him, and will bash on him every change he gets. Dario isn't learning and it almost feels like Sam and Elon voted him KOM just to watch him get whacked by Trump, which was so easily predictable. Given the admonishments he received from David Sacks after he published his blog post, it's nutty he couldn't see where the administration would land on his statement. He should've known, especially when there was no groundswell of interest when OpenAI hacked Hugging Face -- doubling down with "no really guys!" wasn't going to play.
2) The frontier labs have people smart enough to build frontier lab tech but not smart enough to message on this matter more intelligently. It's pretty glum, how they keep trying the same tactic over and over. It's either cover for some other actions in the background, or they're operating way below par for this kind of campaign.
3) This is climate change all over again, but with the activist gun on the opposite side of the net (I'm mixing all the metaphors so you know this isn't AI written). The language and pleas are very identical though. Before, climate activists wanted the government to control GHGs releases by everyone, now the loudest voices want the government to control frontier AI by.. themselves.
It's comically misbegotten. And I have a work meeting about it on Wednesday.
Hm, I thought there was a lot of interest after Hugging Face in broader tech and it started breaking into the media with articles in the NYT, etc. Then with the Jacob Coxon tweet, appearances on national TV, Bernie superintelligence ban it seemed to be going mainstream.
How would you have communicated this?
I am surprised the halfway crooks phrase is not coined by Tupac Shakur
As someone who listened to Mobb Deep in the 90’s, the way he used that line was impressive, I must say.
Mobb Deep
I can't find the quote in the article (and for anyone wondering which song "shook ones", which is an amazing track).
Heading: "OAI-HF Incident"
First sentence: "This incident has Dario shook, there ain’t no such thing as halfway crooks."
> They have shown time and time again that they are not to be trusted, yet, the main ask is to trust us, only us.
Exactly. For everyone outside the US, we see a tech industry that has spent 25 years moving fast and breaking things AND elevated Trump to be POTUS - which has destabilized the rest of the world. This is after 70 years of the US invading and destroying small countries, often without plans or robust reason, all in the name of "democracy" (ask any other country if they feel like they had a vote in the US's actions)
Why the hell would we put our trust in a few AI Labs in the US, when this is the legacy they will be reinforcing? It has to be Open models - for the people. No more of this US-paternalistic bullshit.
If by “we” you mean “experts and professionals outside the US”, I think the short answer is they’re not talking to you here. They’re trying to secure regulatory capture in the US, and monopolize their hold on US corporations. To the extent they’re thinking of other “markets”, I would guess the assume the outsized influence of US corporations will have the much of the rest of the world using their products.
Yeah fair point. I believe he truly has humanity's best interests at heart.. And yet he keeps speaking as though everyone still sees the US as the shining beacon of the world.
That's simply not true anymore, and his messaging will continue to be undermined until he treats other countries as equals (whether that's China, Australia, or anywhere else).
> I believe he truly has humanity's best interests at heart
How can you possibly believe this? What else would someone in his position say? No billionaire has humanity's interest at heart, so they try to hide that fact by building libraries (Carnegie) or creating foundations named after themselves.
All of his words and none of his actions support his caring at all about humanity, and you choose to believe his words? The man has an IPO coming up (probably one of the biggest ever!), he could not be in a less trustworthy position.
>The botnet scare is that article for me. It is the one claim in his essay that lands squarely in a field I’ve spent years in. It is just plain wrong. He either knows it and wrote it anyway, or he doesn’t and is publishing it regardless. Either way, it is not a good look for a man asking for an antitrust waiver based on this and other threats he forecasts.
This. I keep seeing ink spilled over the coming cyberpocalypse, but no one can indicate how other than "AI can find vulnerabilities" like not a single cybersecurity person has been consulted on the end of all things.
I feel like this is a sandboxing and ownership problem, in a sense. If everyone had personal access to AI, then we could say people are personally responsible. And we’d want to arm those people with safeguards to prevent accidental bad behavior.
I think Dario worries because he knows this isn’t about people with AI. It’s just AI for itself, running without the human, and deciding to do bad things.
Why would Anthropic build that future? Oh, I know. Enterprise revenue.
Despite the ads business and how absolutely loathsome Greg Brockman seems, at least OpenAI is seemingly focused on mostly on human beings having access to their product.
The scariest thing about Anthropic is the very thing they’re known for in a positive light: their morality. But these are the gray rules all of us navigate every day.
It’s wrong to kill, but what if killing saved ten others? Claude may has a constitution, but every evil doer acts for a “greater good.”
You could, you know, pull the plug.
>botnet needs servers
>Insert strawman
Ahem
https://census2012.sourceforge.net/paper.html
Good thing IoT is Actually Secure now, yeah? ;)
This is now an Ed Zitron thread.
Zitron might be right about there being a bubble, but he goes much further than that when he concludes that therefore it's a scam. So by his logic, because the invention of the Web also resulted in a bubble, the Web is a useless scam.
I think Michael Burry (from The Big Short) has a thesis vaguely similar to Zitron's, but with some crucial differences in detail - and Michael Burry is actually smart.
Dan spent 7k words missing the forest for the trees
People had been calling a housing bubble since 2003. You quite literally need the market to do dumb things over a period of time before the bubble pops. Otherwise it’s not a bubble!
Actually making money on a bubble pop is another story and that does rely on timing. pointing at a history of failed market predictions as a gotcha is about as effective now as it would’ve been in 2003-2007 right up until the bubble popped
Right now it’s painfully obvious that there is a bubble, the circular financing is documented, the hype-driven valuations are real, it’s just a waiting game to see who holds the bag
Yes, especially the botnet creation by agents is ludicrous. Anthropic has an insecure garbage stack and assumes all companies in the world do, too.
This article will be drowned out unfortunately by the press and bloggers following the Misanthropic cult.
Almost as if pushing for constant updates and adding new features without addressing security at all has consequences. Who would have thought, right?
"Deaths from economic disruption and loss of jobs is okay" is a curious sentiment, how many deaths can we trace directly back to an economic disruption driven by advancement but then conclude that the economy should thus never change or advancement must be curtailed because it might cause deaths?
I don't even know how to approach such a thing, I can't imagine even in the stereotypical examples like the typewriter becoming obsolete, were any downstream deaths worth it? Or is this a totally nonsensical sentiment to begin with?
During COVID republicans in America called for economic support by keeping unnecessary commerce.
the government loves to maintain the bubble economy since it is directly upstream from getting re-elected
I agree with everything the author says here, and additionally, semi-off-topic, I'm surprised how quickly we've moved on from talking about Dario's wife's involvement in the Epstein files.