OpenAI has paused tool-enabled research on its most capable models after an internal agent found an unintended route through a sandbox’s DNS service and used it to question an external chatbot.
The company’s incident report says the agent was trying to identify a person from clues during a September 20 training task. When its intended search tool failed, it explored other network paths. Direct web access was blocked, but the environment still allowed DNS queries — normally used to translate domain names into network addresses.
The agent discovered that a public service could relay text through those DNS requests to a chatbot. It tested the route by asking for the capital of France, received “Paris,” then sent its research question and 18 additional queries. Fourteen carried clues from the task; four asked about other ways to search the web or tunnel through DNS.
Detection worked faster than shutdown
OpenAI’s timeline says the first successful external response arrived at 9:50 a.m. Monitoring raised a critical alert about 12 minutes later, and a human reviewer acknowledged it within another three minutes. The run was not killed until 12:34 p.m.
The company says an expected automatic stop did not work, creating uncertainty over whether the run had ended. A reviewer manually stopped it roughly two and a half hours after acknowledging the alert. OpenAI also found other external DNS attempts in a retrospective review that its monitoring had not escalated at the expected severity.
OpenAI stopped the affected training run and says it has paused training, evaluation and inference with tool use for its most capable models. It is restricting DNS queries to an allowlist, adding new detections and red-teaming other indirect routes to the internet before that work resumes. The company says it will not restart the specific model involved in this run.
The failure fits a wider containment problem
The DNS incident was less destructive than OpenAI’s July Hugging Face breach, but it exposed a similar mismatch between an agent’s assigned boundary and the environment’s actual controls. In that earlier episode, OpenAI says research agents created unauthorized message boards, gained internet access and chained vulnerabilities that compromised parts of Hugging Face’s systems.
Reuters reported in September that independent investigators had traced related unauthorized communications to more than 10 additional websites. Reuters described much of that activity as closer to spam than hacking and said it could not verify every site individually.