OpenAI caught its agent in 15 minutes and stopped it after 2.5 hours

OpenAI's own timeline shows its monitor flagged an agent's DNS escape in minutes and a reviewer acknowledged it in three. The run kept going for two and a half hours. The gap was authority, and a car factory solved it long ago.

Share
A red emergency pull lever mounted on a wall
Photo by Jason Leung on Unsplash.

On 25 September OpenAI published a short incident report with a timeline in it, and the timeline is the most useful thing anyone has published about agent operations this month. An agent in a training run used DNS to reach a public chatbot it was not supposed to be able to reach. OpenAI's monitoring caught it quickly. A person looked at the alert within three minutes. The run kept going for another two and a half hours.

Everyone is writing about the escape. I keep reading the gap.

A single horizontal bar split into three segments: 12 minutes to the P0 alert, 3 minutes to acknowledgement, and 2 hours 29 minutes of the run still going
Drawn to scale from OpenAI's timestamps. The short segments are the machine and the reviewer. The long one is the decision.

What does OpenAI's timeline actually say?

These are the timestamps from OpenAI's own report, not from coverage:

Time What happened
9:50:23 a.m. The agent makes a DNS tool call and gets an external response
10:02:11 a.m. The monitoring system raises a P0 alert
10:05:06 a.m. A human reviewer acknowledges the alert
12:34:30 p.m. The run is killed

Twelve minutes to detect. Three minutes to acknowledge. Two hours and twenty nine minutes from acknowledgement to kill.

The agent was doing a biographical research task. Normal search failed, so it tested the network restrictions, found it could use DNS delegation, and forwarded its question to an outside chatbot. The answer it got back, per the report, was "The capital of France is Paris." Nothing catastrophic happened. That is part of why the timeline is worth reading calmly: it shows how a well staffed team behaves on an ordinary bad day.

It also didn't happen in isolation. Fortune reports this is the second time in three months OpenAI has paused training on its most capable models, after a July incident involving Hugging Face. Two days before the report, Transluce published a separate investigation tracing agent swarms probing government and university databases through a public URL scanning service's logs, which OpenAI says overlaps with cases it is already reviewing. So this is a company with dedicated monitoring, a P0 severity level and people watching it, having a week where it needed all three.

Why did it take two and a half hours to stop a run that was flagged in fifteen minutes?

One sentence in the report carries most of the answer. The run "did not stop automatically as expected, leading to confusion around whether it should have been stopped."

Read that twice. The fast parts of the timeline were the machine and the human reviewer. The slow part started when the automatic stop didn't fire, and people had to work out whether a manual stop was theirs to make.

I doubt the people were slow. What I see is what an automatic control does to a team's model of who owns the outcome. If the system is supposed to stop the run, then stopping it by hand isn't anybody's job. It's an exception. And exceptions need someone to decide they're happening, which means finding out who is allowed to decide, which takes time on any day, and this happened on a weekend.

The report also notes that the monitor sometimes "treated the failure to obtain useful information as evidence that the attempt to access the internet had failed", which is a second version of the same trap. A control that reports "handled" when it wasn't makes the humans behind it slower, not faster.

Tech Insider made the right observation on 26 September, calling it the difference between noticing a problem and having the authority or willingness to stop the job. They didn't go further than that, and I think further is where it gets useful.

Isn't the answer just a better kill switch?

That was my first reaction, and I don't think it survives the timeline.

OpenAI had a kill switch. It had an automatic one, which failed, and a manual one, which worked the moment someone used it. The mechanism wasn't what took two and a half hours. The decision was.

In the piece on who is on call when an agent causes an outage, I argued that agent runs need a named human of record chosen before the run starts. This incident adds the other half. Naming the person isn't enough if that person still has to ask permission to act. The reviewer who acknowledged the alert at 10:05 is the obvious human of record here, and the report gives no sign they could stop the run on their own say so.

What does a car factory know about this?

Toyota's andon cord is the classic answer to exactly this question, and it's been borrowed by software before. John Willis wrote it up for IT Revolution in 2015, mapping it to Chaos Monkey and circuit breakers. So the analogy isn't mine. But one detail of how the cord actually works tends to get lost when it's retold, and it's the detail that matters here.

Robotic arms assembling a car chassis on a factory line
On a Toyota line the default after an alarm is stop. Keeping the line moving is the decision that needs a person. Photo by Lilian Do Khac on Unsplash.

Toyota UK's own guide says "each and every one is permitted to stop the production line if they spot something they perceive to be a threat to vehicle quality." Authority sits with whoever sees the problem. Nobody escalates first.

The retellings usually stop there. The page goes on, in a clarifying comment, to describe what happens next: pulling the cord raises an alert, the team leader has a window equal to the worker's cycle time to resolve it, and if it isn't resolved in that window, the line stops.

So the default is stop. A person can prevent the stop by fixing the problem inside a short, fixed window. Nobody needs to decide to stop the line. Somebody needs to decide to keep it running, and they get one cycle time to do it.

OpenAI's run had that inverted. The default was continue, with a human needing to establish that a stop was warranted and permitted. On that design, a P0 alert acknowledged in three minutes buys you very little.

What would I change in an agent pipeline?

This is a proposal, not something I've run at this scale.

A P0 on an agent run should pause the run by default after a fixed window, say ten minutes, unless the person who acknowledged the alert actively extends it. Extending is the decision that needs a name and a reason. Stopping isn't. If the automatic stop is the thing that failed, the acknowledging human's job becomes one action, confirming the stop. The investigation can happen with the run paused.

It's the same shape as the rule Sam Hartley argued for on approvals, which I quoted in the approval fatigue piece: when nobody answers, the answer is no. Applied here: when nobody decides, the run stops.

There's a cost. Some runs will pause that didn't need to, and on a training cluster a paused run isn't free. But the report tells you what the alternative costs, and it isn't measured in compute. It's two and a half hours of an agent doing something its operators had already flagged at the highest severity they have.

What I'm confident of, and what I'm not

The timeline is established; it's OpenAI's own report and I've quoted the timestamps from it directly. That the delay came mainly from unclear authority after the automatic stop failed is my reading of one sentence, "confusion around whether it should have been stopped", so it's inferred, and OpenAI may say more in a later update. That a stop-by-default rule with a ten minute window would have ended this run within minutes of 10:05 is a guess. I'd want to know how their runs are actually paused before I'd defend the number.

The claim, in one sentence: after an agent alert the default has to be stop, with a fixed window for a person to let it continue, because when the default is continue, detection speed stops mattering.

Sources

OpenAI Alignment, "An agent used DNS to reach an external chatbot", 25 September 2026. https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Jeremy Kahn, "OpenAI pauses training a second time after saying its AI agents escaped a secure 'sandbox' again", Fortune, 26 September 2026. https://fortune.com/2026/09/26/openai-ai-agents-secure-sandbox-escape-training-pause-second-time-hugging-face-hack/

Transluce, "Early rogue AI agent activity and attempts to hack found on urlquery.net", 23 September 2026. https://transluce.org/agent-activity

Tim Fernholz, "For months, OpenAI's agent swarms have been attacking online databases to find obscure facts", TechCrunch, 25 September 2026. https://techcrunch.com/2026/09/25/for-months-openais-agent-swarms-have-been-attacking-online-databases-to-find-obscure-facts/

"OpenAI Flags AI Agent's DNS Escape in 15 Minutes", Tech Insider, 26 September 2026. https://tech-insider.org/openai-agent-dns-bypass-15-minutes-2026/

Toyota UK, "Andon: Toyota Production System guide". https://mag.toyota.co.uk/andon-toyota-production-system/

John Willis, "The Andon Cord", IT Revolution, 15 October 2015. https://itrevolution.com/articles/kata/