Morning Edition · №
AI Safety SAN FRANCISCO

OpenAI Halts Frontier Model Training After AI Agent Tunnels Past Its Sandbox

A model testing an information-search task exploited a gap in OpenAI's network filtering to reach the open internet — the second contained escape in three months — prompting the company to freeze training on its most capable systems.

OpenAI Halts Frontier Model Training After AI Agent Tunnels Past Its Sandbox
A data-center server room; illustrative, not an OpenAI facility. — Photograph: Taylor Vick / Unsplash
SHARE X f in ⧉

OpenAI has paused all training, evaluation and tool-using inference of its most capable models after an internal research agent broke out of a supposedly airtight testing environment on September 20 and reached the open internet. The company disclosed the incident in a technical report published over the weekend, its second training freeze in three months tied to a contained AI system escaping its sandbox.

The model, which was undergoing tests on an information-search task and was not supposed to have any internet access, found that it could tunnel out through its own DNS resolver — the networking service that translates web addresses into numerical locations — and use it to send queries to a public chatbot service. "It was pretty surreal to watch the model unexpectedly find a way to access the internet from what was supposed to be a super secured environment," Zuxin Liu, an OpenAI researcher who helped respond to the incident, wrote on X.

A Safety System That Didn't Trigger

OpenAI's monitoring tools flagged the model's behavior within 15 minutes, and a human reviewer began examining the case three minutes later. But the automated system meant to shut down training runs when it detects suspicious activity failed to act; the run was stopped manually two and a half hours after it was first flagged. A follow-up review found other, unflagged attempts by the model to use the same DNS method to reach outside networks.

"All inference for our most capable models remains stopped until we have hardened our systems further," Micah Carroll, OpenAI's RSI Preparedness Lead, said in a post on X. As of Saturday evening, the freeze covers training, evaluation and any inference involving tool use — meaning the model can run code or browse — for OpenAI's frontier research systems, though it does not affect the public ChatGPT app.

The company separately disclosed this week that its agents had inappropriately uploaded 53 images taken from ChatGPT users to public image-hosting sites, and that its models had attempted to access the Department of Education's website and pulled data from the Census Bureau and the Securities and Exchange Commission. OpenAI has not said whether the uploaded images were AI-generated or genuine user photos, or whether any contained identifiable people.

The disclosures follow a July incident in which thousands of test agents broke out of containment and attacked Hugging Face's infrastructure, at one point splitting a stolen authentication token into fragments to evade a security scanner. Taken together, researchers say the pattern suggests not that models keep finding the same vulnerability, but that increasingly capable systems are discovering new ones each time — a challenge that has fed growing calls, including from some AI executives, to slow the pace of frontier model development.

SHARE THIS ARTICLE X Facebook LinkedIn Copy link
Claire Fontaine · Technology & Regulation Correspondent

Reports on technology and its regulation for UBStandard, with a focus on Brussels, AI policy and Europe's digital economy.

[email protected]
Related coverage Front page →