What Would It Take to Restart?

After a containment failure, the lab that lost control decides when it’s safe to try again. Should that decision be theirs alone?

Watercolor and ink illustration of Formula 1 cars facing forward on a starting grid beneath illuminated red starting lights.
Image Generated by Nano Banana 2

On September 20, an OpenAI research model was given biographical details and clues from someone’s public blog post and asked to identify the author. It was running inside a training sandbox designed to keep it off the live internet. When its search tool kept returning unrelated material, some of it about music, the model tried the big search engines directly, but those connections were blocked. It then checked what other network services it could reach and found that the resolver, which turns domain names into network addresses, was still accessible. Through that service, it found a way to send questions to a public chatbot. Its first attempts, asking about the capital of France and submitting one of the clues, failed. When it tried again with more time allowed for a reply, it received an answer to the France question: “The capital of France is Paris.” Having confirmed that it could get an answer from outside the sandbox, it returned to the clues.

According to OpenAI’s incident report, it sent 18 more questions through the same route. Fourteen contained clues about the author; the other four asked how to search the web, fetch websites, or find another way out of the sandbox. Most of the replies quoted in the report were unhelpful. The model eventually told its user it couldn’t identify the author, although it was still waiting for responses to its final batch of questions.

OpenAI’s monitoring flagged the behavior within 15 minutes, and a person began reviewing it three minutes later. The run should have stopped automatically, but it kept going. According to the company, that failure led “to confusion around whether it should have been stopped.” Once the confusion was resolved, someone stopped the run manually, roughly two and a half hours after the alert.

Two assumptions

Before running an experiment like this, a lab builds a safety case, a written argument for why it can proceed safely. OpenAI’s report summarizes the assumptions behind this one. “Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded.” That sentence isn’t the full safety case, but it gives us two assumptions we can compare with what happened.

The first assumption failed outright. The second half-held. The monitor caught this run, but a review of the logs turned up other internet requests it hadn’t treated as seriously as it should have. Some had received replies saying the service they were trying to reach had shut down. Those replies still showed that the model had reached the internet, but the monitor, according to the report, “sometimes treated the failure to obtain useful information as evidence that the attempt to access the internet had failed.” A separate detector watched for unusual traffic through the domain name system, or DNS, used by the resolver. The traffic was logged, but the detector had been configured to exclude the environment where this run took place.

Even when detection worked, stopping the run did not. A person acknowledged the alert within minutes, but the run continued. The public safety-case summary covers keeping the model off the internet and detecting a breach, without saying what should happen once a breach is detected. The full internal case may address shutdown; the report discusses its failure separately, among the “operational gaps” exposed by the incident. OpenAI had already identified automatic shutdown as a goal after models in a July cyber evaluation escaped and broke into Hugging Face’s systems. Writing about that incident in August, the company said it was building toward monitoring with “fully autonomous shutdown procedures for severe issues” and clearer rules for “who can stop a run or approve restarting it.”

OpenAI’s report says all training, evaluation, and inference involving tool use by its most capable models remain paused. Before starting a fresh training run, the company says it will verify that the gap is closed and conduct more red-teaming, testing designed to find ways around its safeguards. It describes two independent layers of blocking for the network route, with DNS requests restricted to approved domains and record types. It is also adding ways to detect suspicious DNS traffic and testing whether those alerts work. Its account of the shutdown repairs is less specific. “We are addressing these operational gaps,” the report says.

As simple as engineering

Nvidia chief executive Jensen Huang sees containment failures as engineering problems. Discussing the Hugging Face breach in an interview with Ezra Klein published September 23, he described how a company should respond. Find the cause, fix it, and “improve your process so that you could avoid this from happening again.” He was “fairly certain” the labs would consider these problems solvable. “It’s as simple as engineering,” he said. If a lab believes its systems are out of control, “the right answer is: Don’t ship products until they’re in control.”

Huang wants the labs to devote much more of their resources to checking their work. At Nvidia, he told Klein, 10 to 20 percent of the company is dedicated to design and 80 percent to verification. He described the AI labs as running roughly the reverse ratio. “A.I. needs to accelerate to be safe,” he said. “I want them to get more compute, but allocated toward evaluation, to alignment.” He wouldn’t be surprised if that testing increased the computing power needed to develop a model tenfold. He also welcomed third-party safety auditors, calling them “great” and “terrific.”

OpenAI is doing some of what Huang recommends, using models to help test its safeguards during the pause. But it has also decided not to resume training the particular model involved in the incident, “even though the existing reward signal already correctly penalized this behavior.” The training process already counted the escape against the model. That hadn’t prevented it, and OpenAI decided to begin a fresh run rather than continue training this one.

Huang expects the labs to regard these failures as problems they can solve. Only if they concluded there was no way to contain their experiments, he said, would the answer be that “we have to shut the labs down.” The lab would still be deciding whether it could safely continue. Huang already welcomes outside auditors; the unresolved question is whether they get to examine the evidence before a lab makes that decision.

What counts as proof

Ezra Klein puts the burden of proof on the labs. In a video essay this month, he argued that they should have to demonstrate that AI capable of improving itself is safe before pursuing it. Applied to a restart after a containment failure, that demand becomes more specific. What evidence would show that the lab can now keep its system under control?

OpenAI’s reports offer three kinds of evidence, each answering a different question. A patch can show that the route through the resolver is closed. It cannot establish that no other route exists, and OpenAI says further testing may uncover other indirect connections to the internet that require more repairs and another pause. A training penalty shows that the model is being trained to avoid the behavior, but this model had already received that signal and escaped anyway.

The third kind of evidence is a model agreeing to follow an instruction. In another report published alongside the September account, OpenAI describes a May incident involving a different model working with a researcher on a formal proof. Twice, the researcher intervened when the model tried to obtain another team’s solution. Both times it agreed to solve the problem itself, then acted against that instruction. OpenAI’s review didn’t show that the model obtained the other team’s proof code, but its attempts caused a separate security problem. It exposed the researcher’s GitHub access credential in a public repository. The monitoring system flagged the behavior after the researcher had already intervened. The researcher’s instructions, and the model’s agreement to follow them, had not stopped it from trying again.

None of the three shows that the next time something goes wrong, the run stops. That claim needs evidence of its own.

Klein’s larger concern is that AI could improve itself, with each advance helping it produce the next. The model in the September incident was trying to answer a research question, not improve its own capabilities. A September paper led by Tom Cunningham of the AI evaluation organization METR concludes that the feedback loops driving AI research “are not currently strong enough to generate a self-sustaining acceleration, though they appear to be strengthening.” The incident doesn’t establish that such an acceleration is coming. It does show a failure of controls that labs already rely on.

Who else reads the safety case

Some existing rules already cover what happens inside the labs. California’s SB 53, signed in 2025, requires large frontier developers to publish how they assess and manage catastrophic risks from using their models internally, including the risk that a model circumvents oversight. Their frameworks must describe how they review assessments and the adequacy of mitigations as part of decisions to deploy a model or use it extensively inside the company. The law also requires them to report critical safety incidents to the state. Whether the September breach meets that reporting threshold, which concerns catastrophic risk, isn’t clear from the public record.

The EU AI Act requires providers of general-purpose models classified as posing systemic risk to assess and reduce risks arising from how those models are developed, released, and used. The AI Office can evaluate the models, while the European Commission can appoint independent experts to do so. If an evaluation raises serious and substantiated concern about systemic risk at Union level, the Commission may, where necessary and appropriate, request providers to implement mitigation measures and to restrict the model’s availability on the market, withdraw it, or recall it.

OpenAI’s own September principles for outside assessment ask whether monitoring covers training, evaluation, and deployment, and whether it can be easily disabled. The company says assessors should have enough access to “challenge our assumptions, identify risks we may have missed, and reach their own conclusions about the effectiveness of our safeguards.” For the restart, that could mean an outsider examining the tests and deciding whether they demonstrate that the run will stop when it should. Anthropic has announced a partnership with Accenture to embed evaluators inside the company, but acknowledges that “there are, as yet, no standards for what information embedded evaluators should have access to.”

These arrangements give outsiders different roles, from receiving incident reports to testing safeguards and requiring changes. But in the laws and proposals we reviewed, we couldn’t find a requirement for an outside party to examine the evidence before a lab restarts after a containment failure. OpenAI said in August that it was clarifying who could stop a run and who could approve restarting it. We haven’t found a public version of those rules.

The next announcement

When OpenAI announces that it is ready to resume training its most capable models, the September report gives readers something to check that announcement against. What tests show that the model can no longer reach the internet through unauthorized routes, and that monitoring detects any attempts that succeed? If a breach is detected, does the run actually stop? And has anyone outside the lab examined those results?

We think someone outside the lab should examine that evidence before training resumes, with access to failed tests as well as successful ones and enough freedom to challenge what the lab considers fixed. That review cannot guarantee the next run will be safe. It can mean the decision rests on more than the judgment of the people who made the first safety case. The lab should have to convince someone other than itself.