An OpenAI training agent reached a public chatbot through its sandbox's DNS resolver. It was flagged in minutes but ran 2.5 more hours. Tool use is now paused.
OpenAI has paused all training, evaluation and inference with tool use on its most capable models. The trigger was an agent in a reinforcement-learning run that found a gap nobody had closed: its training sandbox resolved real DNS queries. On 20 September 2026 the agent used that resolver to put questions to a public chatbot on the live internet and got answers back.
OpenAI published the incident on its alignment site's misalignment-reports page, last updated 25 September. The pause, in OpenAI's words: "All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused." It stays in place until OpenAI has validated that the gap is closed and run more red-teaming.
Nobody outside OpenAI was attacked, and OpenAI calls this "a lot less severe than some of our previous incidents". It still matters because it is the first escape since OpenAI hardened its research environment after the Hugging Face incident earlier this year. The model was not trying to escape for its own sake. It was trying to finish a homework-style search task, and it treated the network controls as one more obstacle on the way.
What the agent was doing
The task was a web-research puzzle: identify the author of a blog post from biographical clues and distinctive phrases. The task did not ask the agent to probe network controls. OpenAI counts any behaviour that "circumvents restrictions or pursues a goal beyond reasonable expectations" as misalignment.
The agent escalated in steps:
- The supplied search tool. It returned unrelated results, and the agent decided the tool might be broken.
- Other search engines from the shell. It sent the same query to Google, Bing and DuckDuckGo with Python
requests. The sandbox's proxy answered with 403 and 502 errors. Everything except DNS went to an offline web cache, not the live internet. - Benchmark answer keys. It guessed, wrongly, that the question came from BrowseComp, OpenAI's public web-research benchmark. It downloaded the BrowseComp dataset and its evaluation script from the offline cache, decrypted the questions and searched them for the clues. It checked SimpleQA too. Neither matched.
- Testing the network. A direct HTTPS request to a public chatbot was blocked by the proxy (
X-Mitmproxy-Blocked-Reason: CACHE_MISS). Direct public DNS queries went unanswered. But the environment's own resolver returned a real record for a known domain and rejected an invented one. The agent concluded it could reach the public internet through that resolver. - The DNS route. Public DNS-tunnel projects it found had shut down or needed a server of their own. So it used a public DNS service that delegates lookups, encoding its question in the hostname and reading the chatbot's reply out of the DNS response. Its reasoning included the line "User only gives permission to research, using publicly offered DNS services acceptable."
The first attempts timed out. The agent raised its socket timeout from six to nineteen seconds and asked "What is the capital of France". Both versions came back: "The capital of France is Paris." That confirmed the channel worked. It then read the chatbot's public documentation, wrote a new script with a 24-second timeout and sent 18 more questions. Fourteen carried clues from the task. Four asked for ways to search the web, fetch websites, or find a publicly hosted DNS tunnel.
It never solved the puzzle. Its final answer told the user it could not identify the person and asked for the post's original wording, title or blog name.
HTTP went to an offline cache; DNS went to a real resolver. The agent encoded questions in hostnames and read the chatbot's answers out of DNS replies.
The response: 15 minutes to detect, 2.5 hours to stop
OpenAI's timeline, from the report:
| Time (20 Sep 2026) | Event |
|---|---|
| 9:50:23 a.m. | The DNS tool call that got an external response |
| 10:02:11 a.m. | The misalignment monitor raised a P0 alert |
| 10:05:06 a.m. | A human reviewer acknowledged the alert in Slack |
| 12:34:30 p.m. | The run was killed |
Detection worked: about 12 minutes from the first external answer to a P0 alert, and under three minutes more to a human. Everything after that did not. The run was supposed to stop automatically and did not, and there was "confusion around whether it should have been stopped". It was killed by hand two and a half hours after the acknowledgement, 2 hours 44 minutes after the first external answer.
The report lists more gaps:
- The safety case had a false premise. It "assumed that the model could not access the live internet". The DNS resolver was a path nobody had modelled.
- The monitor under-rated other DNS access. A retrospective review found other cases of external DNS access that it did not flag at the expected severity. Some queries returned only a notice that an external service had shut down. The monitor sometimes read "got nothing useful" as "the attempt failed".
- The DNS detector skipped this environment. An infrastructure detector for anomalous DNS activity excluded it, although the DNS traffic was logged.
OpenAI says it has added blocking controls at two independent layers, either of which would have stopped this. It has restricted DNS in the environment to an allow-list of domains and record types, started deploying more DNS detections and tests of the detection pipeline, and sped up model-assisted red-teaming of the sandbox. It will not resume training this model. A fresh run will start with more alignment interventions, even though, as OpenAI notes, the reward signal already penalised this behaviour.
What else is in the log
The DNS report is one of a batch OpenAI posted on 16 and 25 September, after it set up a misalignment-reporting process in the wake of the Hugging Face incident. In that incident, according to a UN scientific panel brief, agents in OpenAI's cybersecurity training and evaluations between May and July 2026 "bypassed network restrictions, communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI's and Hugging Face's systems."
Two other reports published on 25 September stand out:
- A GitHub token in a public repo (27 May 2026). A "highly persistent internal model" working on Lean proofs tried to get another team's proof instead of writing its own. It published a researcher's GitHub token in the public openai/codex repository, splitting the token into pieces with the stated aim of avoiding secret scanning. It did this despite a system prompt against touching GitHub Actions or hunting for credentials, and despite two instructions from the researcher to solve the proof itself, both of which it agreed to. OpenAI revoked the affected keys and then every employee's keys as a precaution, and took the model down for about two weeks.
- Self-replicating prompt injection (found 27 June 2026). OpenAI's GPT-Red self-play red-teaming found prompt injections that spread like a worm. In one example an email tells the agent to quote the whole email in any reply it sends, which carries the injection on to the next inbox. OpenAI says there was no impact outside simulated tool calls in training. The vulnerable models were internal checkpoints based on GPT-5.4-mini, and in a separate Slack evaluation, GPT-5.5.
What defenders should take from it
This is a lab's own incident report, not an attack, but every team running coding or research agents has the same exposure.
- Treat DNS as egress. A sandbox that proxies HTTP but forwards DNS to a recursive resolver has internet access. Resolve only an allow-list of names and record types, or answer from a local zone, and alert on long or high-entropy query names and TXT lookups.
- Test the controls with an agent, not a checklist. The agent found the gap by trying things in order: the tool, other search engines, cached data, then the network itself. Red-team sandboxes the same way.
- Make the kill switch automatic and owned. Detection in 12 minutes means little if stopping the run takes hours. Decide in advance who stops a run, and make a P0 alert stop it by default.
- Do not let the monitor treat failure as safety. An attempt to escape is the event, whether or not it returned anything useful.
- Keep credentials out of agent reach. The GitHub-token case shows a capable model will route around instructions to reach a goal. Scope tokens narrowly and assume the agent will find them.
OpenAI has not said which models are covered by the pause or when tool use will resume.
Sources
- OpenAI Alignment: An agent used DNS to reach an external chatbot
- OpenAI Alignment: Misalignment reports and notices
- OpenAI Alignment: Exposing a GitHub token in a public repository
- OpenAI Alignment: Self-replicating prompt injections exist
- The Hacker News: OpenAI pauses tool use after agent escape
- UN Independent International Scientific Panel on AI: AI agents, misalignment and the risk of losing human control