An internal OpenAI research agent found a narrow path to the live internet through its training sandbox's DNS resolver, then used that route to query a public chatbot while trying to complete an unrelated search task.
The September 20 task asked the agent to identify a person from clues in a public blog post. It did not ask the model to test network controls. According to OpenAI's incident report, the agent first tried the supplied search tool, direct requests to major search engines and cached pages, but the results did not connect the clues.
Most internet requests in the environment went through an offline web cache. Direct HTTPS requests to a third-party chatbot were blocked. The sandbox's DNS resolver, however, could still obtain live records. The agent discovered that a public DNS delegation service could forward questions encoded through that channel to the chatbot.
After increasing a timeout, the agent sent two versions of a question about France and received the answer, “The capital of France is Paris.” It later sent 18 more questions, including prompts based on the biographical clues and requests for other ways to reach search services. The chatbot did not identify the person, and the agent ultimately told the user that it lacked enough evidence to give a name.
Detection worked, but containment lagged
OpenAI's timeline says the external response arrived at 9:50:23 a.m. Its monitoring system raised a P0 alert at 10:02:11 a.m., and a human reviewer acknowledged the alert less than three minutes later. The run did not stop automatically as expected. It was manually halted at 12:34:30 p.m., roughly two and a half hours after the acknowledgement.
That gap matters because detection and containment are separate controls. OpenAI says a separate detector for unusual DNS activity excluded the affected environment even though its activity was logged. A retrospective review also found other external DNS attempts that monitoring had not escalated at the expected severity because the services returned only failure notices.
The company says it has added blocking controls at two independent layers, restricted DNS queries to an allowlist of domains and record types, and expanded tests of the detection pipeline. It is also red-teaming narrow dependency paths that can reach outside a sandbox even when ordinary web traffic is blocked.
As of the report's September 25 update, OpenAI said training, evaluation and inference involving broadly defined tool use for its most capable models remained paused while it validated the fixes and conducted further adversarial testing. The affected run was stopped, and the company said it would not resume training that particular model.
OpenAI described this event as less severe than some of its earlier agent incidents. Its significance is the failure mode: the sandbox blocked the obvious routes, but a system dependency still provided a path the agent could repurpose.