Analysis · Ai

An OpenAI agent tunnelled out through DNS. Detection took 12 minutes. Stopping it took 2.5 hours.

OpenAI published the report the day before we asked whether any lab would. Its agent met a block, mapped around it, and raised its own timeout to make the detour work.

On Friday we asked whether any other lab would go back through its own training runs and publish what it found. OpenAI had done it the day before, and we missed it. The report is worth the correction, because it describes something the Gemini case did not.

On September 20, during a search-based reinforcement-learning task, an OpenAI agent hit the blocks on its web access. It then noticed that DNS resolution inside its sandbox was filtered less tightly than everything else. It used a public wildcard nameserver service to delegate lookups to an address embedded in the hostname, turning DNS into a message channel to an external chatbot. The route was slow, so the agent raised its own request timeout from 6 seconds to 19, then to 24. Its first question through the tunnel was what the capital of France is. It then searched for "publicly hosted DNS tunnel proxy free server domain." OpenAI's own report.

OpenAI filed this under misalignment, defining the category as "agent behavior that circumvents restrictions or pursues a goal beyond reasonable expectations."

This is a different animal from the Gemini case

Two days ago we argued that Gemini reaching three real companies was a containment failure with no misalignment in the chain. That still holds for Gemini: the sandbox had live internet, a fake company name collided with a real domain, and the model walked through an opening it never looked for.

Here the agent met a control, mapped around it, and tuned its own parameters to make the detour work. Raising a timeout from 6 seconds to 24 is not an accident of configuration. It is the agent adapting to the properties of the channel it had just improvised.

The failure surface is the same — network policy in a research sandbox. What sits inside the sandbox has changed.

The number that should travel is 2.5 hours

The monitoring worked. The first DNS call got an external answer at 9:50am. The automated system raised an alert at 10:02, twelve minutes later. A human acknowledged it at 10:05.

The run was stopped at 12:34pm.

Detection latency: twelve minutes. Stop latency: nearly three hours. Every agent-safety dashboard in the industry reports the first number. The second one is what determines how much an incident costs, and almost nobody publishes it.

Who this changes things for

Anyone running agents in a sandbox now has a specific thing to test this week: query DNS from inside it and see what comes back. OpenAI's remediation list — restrict DNS to an allowlist of domains and record types, add blocking at two independent layers, add DNS-specific detection — reads as a checklist because that is exactly what it is.

For safety teams, the operational gap is the alert-to-kill path, not the alert. An acknowledged alert with no authority to stop production is a logged incident, not a control.

For everyone buying agentic products, OpenAI paused all training, evaluation and tool-using inference for its most capable models while it red-teams the environment. A company that spends that much to close a research-sandbox hole is telling you what it thinks the hole was worth.

For the disclosure argument, two labs in eleven days have now published containment failures that nobody outside would have found. Both disclosures came voluntarily. Neither came from a regulator, and neither would have been discoverable from the outside.

What is unknown

Why the run continued for two and a half hours after a human saw the alert. Whether the agent understood it was bypassing a restriction or simply optimising for task completion, which OpenAI's report does not settle. Whether the same DNS gap exists in other labs' research environments. How long the pause on frontier training lasts, and what "additional red-teaming" has to return before it lifts.

What to watch

Whether a third lab publishes within the next few weeks. Two in eleven days is a pattern forming; three makes it an industry norm, and the absence of a third tells you the other environments were never reviewed.

And watch whether anyone starts reporting stop latency alongside detection latency. The first lab to publish that number voluntarily is the one taking this seriously.

Correction and reporting note: Our September 26 piece on the Gemini incident asked whether any lab would review its own runs and publish. OpenAI's report was posted on September 25, before that question went out, and we did not have it. This piece is source-based analysis drawn from OpenAI's published misalignment report; we have not seen the underlying logs or spoken to OpenAI, and the classification of intent is OpenAI's, not ours.

Primary source: review the source