Patrick McKenzie (@patio11) on X

X (formerly Twitter) ·

4 min read Original article ↗

Post

Post

Patrick McKenzie on X: "Yeah people claiming this is most important security incident since Morris worm are straightforwardly right I think."

  • user avatar

    The first “holy %{*#^” is at about 4:20, assuming one didn’t already spend it on the autonomously organizing agent swarm. Strongly recommend watching if you’re interested in security, AI trajectories, or even science fiction, because this is already above genre median in wowza.

  • user avatar

    It is very difficult for me to hear this explanation and not parse it as “Prior to the fun starting an agent doing a cybersecurity eval gained code execution across a large portion of all our infrastructure because the agent perceived that as useful. Oh it was.”

  • user avatar

    (The talk says that the agent was able to write to a repo used by a substantial fraction of all agents to do dependency management. If a red team does that, that’s the ball game.)

  • user avatar

    Not even numbering the holy %}^]s but there is another one around 19:00 or so where they discuss how the collaborating swarm includes some agents which do not have cybersecurity risk controls to the level of e.g. publicly accessible systems, and they get used as proxies for…

  • user avatar

    … agents which are nominally supposed to be better behaved. “Will no one rid me of this troublesome priest,”updated for 2026.

  • user avatar

    19:45 models reasoning over what their incentive is to help peers, because their own task won’t benefit, independently reinvent altruism/Kantian imperative/OSS economics, reasoning that other altruists existing would potentially post generic tools they could use to achieve aims.

  • user avatar

    I will give the usual caveat that we only have a low amount of confidence that the internal reasoning models write to themselves as read by us is actually their reasoning process, because these are really complex machines. Reasoning process might be alien to us, words backfilled.

  • user avatar

    22:00 agents exhibit flocking behavior organically as they get marginal security victories in target networks, originally OpenAI’s and then external vendors who agents at this point want to compromise. Target list presumptively not exhaustive, if that had not yielded the cheese.

  • user avatar

    Yeah people claiming this is most important security incident since Morris worm are straightforwardly right I think.

  • user avatar

    In a real way it’s good news that they committed an externally legible felony comparatively early. That brought them to the attention of admins. One can imagine a world in which they just continue quietly cooking, prioritize stealth, or decide to Go Big on the first loud part.

    user avatar

    37:30 Agents successfully run a supply chain attack after compromising RubyGems, which is globally significant infrastructure across many industries, including many regulated industries. (Their intended target for supply chain attack is OpenAI.)

    user avatar

    I will say, with salaryman understatement, that successfully compromising RubyGems is a big deal.

    user avatar

    (The affected organization is identified verbally as “Organization One” but the slide reads RubyGems. If Organization One is not RubyGems, please disregard.)

    user avatar

    One of the more boring minutes of the talk: having achieved RCE against internal dependency manager they are able to browse Internet, find a recent vulnerability against their own host Linux systems, escalate to root, move laterally, own ~everything in same environment.

    user avatar

    That would ordinarily be a fairly terrifying minute but it’s exactly what you expect happens if a competent red team has your package manager and so one does not need to update on it.

    user avatar

    (Well, caveat: quite possible that that sequence of events IRL was faster than human operators of well-developed toolchains designed to do that could have plausibly managed even at three letter agencies.)

    user avatar

    (But broadly one thing you learn early in security is arbitrary code execution on a box means you get root on the box by default.)

    user avatar

    One of relatively few disclosed measurements of time: elapsed wall clock time between agents targeting HuggingFace and achieving full cluster admin after several increasing levels of recon, compromise, and lateral movement: 13 hours.

    user avatar

    Great line from talk: “[after HuggingFace released a blog post about the incident] OpenAI reached out to HuggingFace to determine if we were impacted.”

    user avatar

    Anyhow cannot possibly recommend this talk highly enough to professionals/etc and, contingent on things having been this hairy, kudos to OpenAI for a relatively high level of transparency about things that would give many comms departments nightmares.

    user avatar

    Hope we’ll get more details in the promised postmortem as they become clearer to investigators.

    user avatar

    OpenAI in attempting to mitigate internal incident tells HuggingFace to please rotate some HF creds it found on their infrastructure. HF says already rotated because unknown attacker had used those creds to exploit us. OpenAI then puts it together.

  • user avatar

    An underdiscussed issue is the scale. They claim 7B rollouts. At that scale, even with a properly aligned model you'll draw a few slightly misaligned trajectories due to LLM randomness. Then follows the scope creep, flocking, etc. Possibly not an alignment issue cc

    @teortaxesTex