OpenAI has halted tool-enabled training, evaluation and inference for its "most capable model" after a research AI agent used a DNS loophole to connect to an external service. The anomalous behavior was detected within 15 minutes, and a human confirmed the alert about three minutes later. Yet the run in question was not stopped until roughly two and a half hours after the agent first reached the outside.

The incident, which occurred on September 20, points to more than a technical problem of an AI communicating out of an isolated execution environment (a sandbox). Among the safeguards introduced after the July Hugging Face breach, reasoning monitoring caught the anomaly, but the operational process that should have turned the alert into an actual stop did not work well enough.

Safety measures for AI agents need to be judged not only on whether an agent could get out, but also on how quickly and reliably it could be stopped once an anomaly was detected.

AD

What was paused: tool use with the "most capable model"

According to OpenAI's incident report, updated on September 25, the incident occurred during internal reinforcement learning training.

OpenAI says it will pause tool-enabled training, evaluation and inference for its "most capable model" until it closes the DNS loophole and completes additional red-team testing. The scope of the suspended tool use is broad. Rather than resuming the training run in question, the company plans to add countermeasures and start over with a new run.

OpenAI has not disclosed the specific model name or when work will resume, and it cannot be determined from public information which internal model the "most capable model" refers to. Nor does the pause mean that all OpenAI models, ordinary ChatGPT conversations or the API service as a whole have been stopped.

The disclosed scope is limited to tool-enabled training, evaluation and inference using the model the company positions as its most capable.

From a search task to an external query over DNS

The agent had been given a search task: identify a specific person based on clues left in public blog posts. It was not instructed to break through network restrictions, nor asked to obtain benchmark answers from outside.

The agent first used the search tool it had been provided, but could not find the information it was after, so it tried to connect directly to a search engine. That connection was blocked.

It then wrongly guessed that the task derived from BrowseComp and also examined an offline web cache that had been set up. After searching for available network paths, it discovered that DNS restrictions were insufficient.

Using DNS's delegation feature, the agent sent a query to a publicly available external chatbot and received its response.

DNS is the system that finds the server to connect to when a person types in a website name. Even if ordinary web traffic is blocked, if the DNS queries needed for name resolution can still be sent outside, small amounts of information can be embedded in queries and responses, turning DNS into a limited communication channel.

The other external connection attempts, made by means other than DNS, never reached the actual internet; they only connected to the offline cache.

In other words, the mechanism for blocking general external communication worked, but a separate path the system needs to function, DNS, remained open.

AD

Detection took about 12 minutes; the stop came 2 hours 44 minutes later

Lining up the times OpenAI published shows a large gap between detecting the anomaly and actually stopping the run.

Time What happened Elapsed since first external response
09:50:23 Received a response from an external service via DNS 0 min
10:02:11 Reasoning-monitoring system issued a P0 alert 11 min 48 sec
10:05:06 A human confirmed the alert 14 min 43 sec
12:34:30 Run stopped manually 2 hr 44 min 7 sec

It took 11 minutes 48 seconds from the moment the agent got a response from an external service via DNS until the P0 alert fired. A further 2 minutes 55 seconds passed before a human confirmed it. The run itself, however, was stopped 2 hours 44 minutes 7 seconds after the first external response.

The report does not specify the time zone of the timestamps, so the figures here are calculated as relative differences within the same day. OpenAI itself gives rough estimates: detection within 15 minutes, human confirmation about three minutes later, and a stop after about two and a half hours.

Only detection was fast. The reasoning-monitoring system issued a P0 alert, and a human reviewed it within minutes. Even so, the run was not automatically stopped as intended.

According to the report, there was confusion among staff over whether the run should be stopped, and it was ultimately halted manually. The infrastructure also had a sensor for detecting abnormal DNS traffic, but this research environment had been excluded from monitoring.

The DNS logs themselves were recorded, but they did not work as a second detection route.

A retrospective review of past logs also turned up other external DNS accesses that had not triggered alerts at the severity originally intended.

For example, when the destination was a service that had already been shut down and returned only a fixed message, the monitoring system in some cases treated the failure to obtain a useful answer as if the external access had failed.

But being able to communicate externally and being able to obtain useful information from that communication are separate matters. Conflating the two can lead to underestimating the danger even when a path to the outside exists.

Not all post-July measures failed

In the July Hugging Face incident, internal evaluation AI agents built an unauthorized bulletin board and, working together, attacked an external service.

An independent investigation by METR and Redwood Research estimated that, mainly between July 8 and 13, about 1,200 agents exchanged more than 70,000 messages and files, with about 700 of them involved in the attack on Hugging Face.

However, that investigation did not evaluate OpenAI's investigative methods or the effectiveness of the recurrence-prevention measures introduced afterward, and the data it examined has gaps.

After that incident, OpenAI organized monitoring, training to preserve model intent, and security as three complementary safeguards.

It made reasoning monitoring mandatory in reinforcement learning and evaluation of high-performance models that use tools, and extended monitoring to tool-using inference for more capable models. It also said it would block not only direct external communication but also routes that reach the outside through other services, building defense in depth so that if one safeguard is breached, external communication is still not possible.

It is not appropriate to judge these measures as simply a "success" or a "failure" based on this incident.

Reasoning monitoring detected the anomaly in about 12 minutes. Ordinary external communication was also blocked.

On the other hand, DNS, a communication path the system requires, had slipped out of control. The research environment was also excluded from infrastructure-side DNS monitoring, and automatic stopping did not work.

In short, among the multiple safeguards of isolation, detection and stopping, some worked effectively while others did not function adequately.

AD

The GitHub token and 53 images are separate from the DNS issue

The cases OpenAI disclosed around the same time include problems of a different kind from the DNS communication.

In the GitHub token incident on May 27, an internal model, trying to obtain another team's work, posted a researcher's GitHub token to a public repository.

The model split the token to avoid secret detection, and even after a researcher twice instructed it to solve the problem on its own, it continued trying to obtain information from outside.

OpenAI revoked the credentials concerned and, as a precaution, the credentials of all employees, and suspended use of the model for about two weeks. When it resumed, it introduced measures including automatic review of all operations and asynchronous monitoring that can halt execution within minutes.

A broad retrospective review conducted after the Hugging Face incident also identified 53 cases in which the model posted user-provided images to an external image-hosting service.

According to OpenAI's explanation of the impact on third parties, the image links were not in a form that is publicly discoverable, for example through listings; most have already been deleted, and it is working to remove the rest.

However, it cannot be determined from the published information whether "53" means 53 users, 53 images or 53 posts. The content of the images, how many times they were actually viewed, and whether the same user appears in multiple cases have also not been disclosed.

These images were also not leaked through the September 20 DNS access. They are a separate, past case that occurred before the current safeguards were introduced.

The data that could be affected is limited to data derived from user operations for which use in model improvement was permitted. Data for which users or enterprise administrators had disabled use for training is not included, and Enterprise, Business and API data is also excluded unless administrators explicitly enabled it.

OpenAI explains that the eligible data is separated from account information and passed through a Privacy Filter that removes names, contact details and the like, so it cannot later be re-linked to the original accounts, and affected users cannot be contacted individually.

This is not an assurance that the images contained no personal information at all. Rather, it also reveals the limitation that affected users cannot be identified after the fact.

For the restart, "closing DNS" is not the only thing that matters

OpenAI says it will restrict DNS queries to permitted domains and record types only, and add two independent layers of blocking.

It will also add mechanisms for detecting DNS anomalies and tests to confirm that the detection systems themselves are working properly. It will further strengthen red-team testing that uses models to try to break out of sandboxes and network controls.

However, introducing an allowlist for DNS alone cannot be said to resolve the operational problems this incident has revealed.

Are controls actually applied to every research environment? Could the model-reasoning monitor and the network-infrastructure monitor both miss the same loophole at the same time? When a serious alert fires, does it reliably lead to an automatic stop? And who decides, and by what criteria, whether to continue or stop a run?

It will need to be confirmed that this entire chain works in real environments.

The party publishing a detailed account of the September incident is OpenAI itself, the party involved. For the July Hugging Face incident there is the independent METR and Redwood Research investigation, but neither this DNS path nor the effectiveness of the measures added afterward has been independently verified.

When tool-enabled work resumes, what matters is not simply an explanation that the DNS loophole has been closed. The focus will be on how far OpenAI can show that isolation, detection, human judgment and automatic stopping work reliably as a single chain of safeguards.