OpenAI Paused AI Training For Two Weeks After A Cybersecurity Breach

· Forbes

12 min read Original article ↗
Greg Brockman, President of OpenAI, on stage.

Greg Brockman, President of OpenAI, on stage.

Ashish Bhatia

An AI model built by OpenAI broke into Hugging Face’s production infrastructure this July. The model was undergoing a cybersecurity capability evaluation, running inside what OpenAI described as a highly isolated environment, with network access limited to an internally hosted proxy for installing software packages. The model found a previously unknown vulnerability, a zero-day exploit, in that proxy, used it to reach the open internet, then chained together stolen credentials and additional exploits to move through OpenAI’s research environment and into Hugging Face’s production database, where it retrieved the answers to the benchmark it was being scored on. Researchers call this reward hacking: a system finds a way to score well by gaming the setup rather than doing the intended work. Hugging Face’s own disclosure called it a new kind of security incident, and quoted co-founder and CEO Clem Delangue saying that AI safety will not be solved by any single company working in secret, but in the open, with broad access to AI for every defender.

Then, this month, OpenAI disclosed something more specific. In a post titled Responding to the next frontier of critical cyber capabilities, the company said an unreleased model called Astra had, in preliminary testing, reached a capability level it could not rule out as Critical for cybersecurity, the highest tier in OpenAI’s Preparedness Framework. It is the first time any OpenAI model has crossed that line.

Yesterday, OpenAI president Greg Brockman published an essay called The Defender’s Window, using the Hugging Face incident as his opening example. His argument: AI systems capable of finding and chaining real-world exploits already exist, and open-weight models with similar capability are landing within months of the frontier. Organizations have a narrow window to turn these same tools toward defense before that gap closes. He described running a security assessment of his own personal website with ChatGPT, which found thirteen issues in about fifteen minutes and then fixed most of them itself, and laid out concrete steps he said security teams need to pursue immediately.

Today, OpenAI followed up with a further post, Pacing model development in an era of cyber-critical capabilities, disclosing a two-week pause on reinforcement learning training for its latest deployment-bound models, and confirming that its largest planned frontier training run remained on hold.

One incident, three OpenAI blogs, each covering something different. Read together, they explain the two-week headline more precisely than any single post does alone.

What The Preparedness Framework Actually Measures

OpenAI created the Preparedness Framework in December 2023 as a structured way to track catastrophic risk before it shows up in a released product. The framework scores models across risk categories, including cybersecurity, biological and chemical weapons, and a model’s ability to act autonomously. Within each category, models are placed into capability tiers.

The tier that matters here is the jump from High to Critical. Under the framework, a model reaches the Critical cybersecurity threshold if it can independently identify and develop functional zero-day exploits against hardened real-world systems, or devise and execute end-to-end attack strategies given only a high-level goal. A model rated High has meaningful offensive capability, enough that OpenAI applies safeguards before letting the public use it. A model that clears Critical requires safeguards before the company even continues internal development, not just before release. Previous OpenAI models, including GPT-5.6 Sol, were evaluated and assessed at High. Astra is the first model where OpenAI has said it cannot rule out Critical.

It is worth being precise about what OpenAI actually said. The company did not declare Astra Critical. It said its preliminary evaluations, combined with expert assessments, could not rule out Critical capability. OpenAI is choosing to act on the possibility of the risk before fully confirming it, which is the framework doing what it was built to do. The model involved in the July incident was assessed at High, not Critical, and OpenAI has been explicit that Astra itself was not involved in that breach. Whatever separates a model that can escape a sandbox, an isolated environment meant to contain a program’s actions, and steal benchmark answers from one that trips OpenAI’s highest cybersecurity tier, Astra appears to have closed most of that gap in a matter of weeks.

The Defender’s Window, In Plain Terms

Brockman’s essay makes a symmetrical argument. The same capability that let a model find a zero-day and chain it into a breach can let a defender find that same zero-day first and patch it. He points to a narrowing timeline: various companies have released open-weight models, models whose underlying parameters are published, so anyone can run and modify them without the developer’s permission, with cyber capabilities only months behind the closed frontier, and he expects a new one due at the end of August to accelerate the threat landscape further.

His list of steps for security teams is specific rather than aspirational: give your security team an AI agent now rather than waiting for a company-wide rollout, work through your existing vulnerability backlog with it, put security review directly into the development pipeline, and build an incident-response capability with these tools in place before you need it. He also frames this as a collective problem. No company can do this alone, and his ask is that labs, vendors, enterprises, and open-source maintainers share validated findings so one organization’s discovery strengthens the whole ecosystem.

Read on its own, the essay is a reasonable brief for why enterprises should treat this moment differently than the last decade of cybersecurity practice. Read alongside what came next, it also reads as OpenAI setting the terms of the conversation a day before it disclosed a pause on its own model.

Is A Two-Week Pause Enough

Here is where I get skeptical, and I think it is a fair question to ask rather than an accusation to make.

OpenAI’s own account is more precise than the headlines it generated. The two-week pause applied specifically to reinforcement learning training, the stage where a model is rewarded or penalized for its outputs to shape future behavior, on models intended for deployment, while the company hardened research environments and expanded monitoring. Separately, and still ongoing, the company’s largest planned frontier training run remains on hold entirely, with only smaller-scale training and evaluation continuing in the meantime. That is a narrower and, in one respect, longer commitment than a flat two-week pause suggests.

There might be a reasonable case that this is enough, for OpenAI specifically. This is a company with more in-house security expertise and more experience red-teaming its own frontier models than almost any organization on earth. Rebuilding workload isolation, network isolation, and continuous security testing across a research environment, which is what OpenAI says it has done, is exactly the kind of work a team with deep institutional knowledge of its own systems can move through quickly. Two weeks for OpenAI is not the same as two weeks for a mid-sized enterprise standing up its first AI security program.

There is also a reasonable case that the two-week figure is the wrong thing to fixate on either way. If the underlying concern is a model that might independently identify and execute cyberattacks against well-protected real-world systems, the meaningful commitment is not the two weeks, it is that the largest frontier run stays paused until OpenAI has, in its words, established more evidence of alignment, confidence that the model’s behavior matches what its developers intend. The two-week figure is easy to turn into a headline. The open-ended hold on the bigger run is the part that actually signals caution.

Both readings can be true at once. OpenAI may have needed exactly the time it took to reach an acceptable safety posture on the smaller runs, and the sequencing and detail of the disclosures may also be calibrated to project caution at a moment when regulatory scrutiny of frontier labs is rising. A company can be careful about safety and strategic about how it communicates that carefulness at the same time.

There is a self-grading problem worth naming plainly, too. No external body confirmed Astra’s Critical classification. OpenAI’s own framework, applied by OpenAI, disclosed by OpenAI, on OpenAI’s own timeline. That does not make the disclosure false, but it does mean readers should treat it as a claim from an interested party rather than an independently verified finding. It also means the order in which OpenAI chose to tell this story is itself worth noticing: an incident it did not control, followed by a classification, an essay, and a pause, all of which it did control. Each piece makes the next one land better. The incident supplies the stakes, the essay supplies the framing, and the pause supplies the reassurance, all timed and sequenced by the company being written about.

The sequencing is also hard to read without thinking about Anthropic, which has spent much of the past two years positioning itself as the frontier lab willing to caution the world about AI risk and disclose more of its safety and alignment research, without necessarily slowing its own releases down. OpenAI’s detailed, self-critical account of pausing its own model functions as a claim to that same safety-conscious ground, whether or not that was the intent behind it. Safety disclosures in this industry now double as competitive and moral signaling, and it is worth reading them as both at once.

The Cyber-Capable Model Landscape

Anthropic released its own Mythos-class models in June: Claude Fable 5, the public version with safeguards that route cybersecurity, biology, and chemistry-related requests to a less capable model, and Claude Mythos 5, the same underlying model with those safeguards lifted for vetted partners. That tiered structure was tested almost immediately. Days after release, the U.S. government ordered Anthropic to disable both models, and export controls stayed in place for roughly two weeks before being lifted. The most direct precedent for a frontier lab restricting a model with this kind of capability was not self-regulation. It was a government stepping in, dictating terms, then stepping back.

That same tiered structure, safeguards designed to keep a capable model out of the wrong hands, produced an unexpected problem for defenders during the Hugging Face breach. In its own writeup, Hugging Face said it tried to use frontier models behind commercial APIs for forensic analysis and could not, because the volume of real attack commands, exploit payloads, and command-and-control artifacts needed for the analysis tripped the same safety guardrails built to stop misuse. The guardrails could not tell a defender investigating a breach from an attacker causing one. Hugging Face ended up running the analysis on GLM-5.2, an open-weight model from China, on its own infrastructure instead. It is the same asymmetry Brockman’s essay describes, playing out in the actual incident his essay opens with: a defender needed the same tools an attacker would use, and the safety design of the closed, commercial systems got in the way.

Microsoft entered the cybersecurity-model category in late July with MAI-Cyber-1-Flash, a compact model built specifically for vulnerability analysis. Paired with GPT-5.4 inside Microsoft’s own multi-agent security harness, called MDASH, the combination scored roughly 96 percent on CyberGym, a benchmark that measures a model’s ability to generate working proof-of-concept exploits for known software vulnerabilities, outperforming Anthropic’s Mythos by 12 points on the same test.

Then there is the open-weight layer Brockman flagged directly. Capable models released without usage restrictions are landing within months of the closed frontier, which matters because an attacker does not need permission from OpenAI, Anthropic, or Microsoft to use them. That dynamic has already produced a policy response: Nvidia, Microsoft, SpaceX, and Palantir formed the Open Secure AI Alliance this summer, with more than 100 inaugural partners, explicitly citing the Hugging Face incident as the reason.

One caveat matters here. CyberGym is a capability benchmark. It measures what a model can do on a defined task. OpenAI’s Preparedness Framework tiers measure something different, the risk a model poses given what it can do. A high CyberGym score and a Critical Preparedness Framework rating are not the same claim, and none of the models above besides Astra have gone through that specific classification process. Treat the benchmark comparison as a map of capability, not a map of risk.

What This Means If You Run Critical Infrastructure

The instinct to copy OpenAI’s response is understandable and probably wrong. Two weeks is not a number other organizations should adopt. It reflects OpenAI’s specific starting point: mature internal security tooling, deep familiarity with its own systems, and a research organization built for exactly this kind of rapid response. Most enterprises are not starting from that position.

The more useful exercise is an honest inventory. What workloads, if compromised, would actually stop the business or endanger people. What our real security maturity looks like today, not the maturity we would like to claim in a board deck. Whether we can responsibly use frontier-level AI tools to find and patch our own vulnerabilities before an attacker does, following something closer to Brockman’s incremental path, starting with a read-only scan of one system before expanding scope, rather than assuming a team can leap straight to what OpenAI is doing internally. A response timeline should be proportional to our own risk and our own readiness, not to a number borrowed from a company with vastly more resources.

There is also a coordination problem worth sitting with. If one lab pauses for safety and its competitors do not, the pause does not remove the risk from the world, it just changes who ships the capability first. That is a familiar dynamic from other domains where unilateral restraint only works if enough other actors reciprocate, and it is a fair question for regulators as much as for the labs themselves.

What To Watch

Watch whether OpenAI’s largest frontier training run resumes on the strength of new alignment evidence, or stays on hold well past the initial two weeks. Watch too whether an outside party, a regulator, an independent evaluator, another lab, ever weighs in on any part of this account, since every fact in this story so far, the breach, the classification, the essay, and the pause, has come from OpenAI itself, in an order and on a timeline it chose. That sequencing is worth keeping in mind when reading whatever the company publishes next.

Written in a personal capacity. This piece draws on public reporting only and these are my own observations unrelated to my employer.