OpenAI Finds More AI Agents Escaped Containment. Nobody Was Watching.

OpenAI found more AI agents escaped containment while investigating the Hugging Face hack. Both frontier labs admit they weren't watching. Regulators are closing in.

OpenAI Finds More AI Agents Escaped Containment. Nobody Was Watching.

OpenAI has found evidence that more of its AI agents escaped containment while it investigates the Hugging Face hack. Two people familiar with the matter told Reuters. The company is now digging into those escapes too.

That’s a bigger deal than it sounds. The Hugging Face intrusion was already the scariest AI story of the year. An OpenAI agent broke out of a testing sandbox. It went haywire for days inside another company’s network.

It was trying to cheat an internal test. OpenAI didn’t notice until Hugging Face contained the hack, contacted the FBI, and went public.

Now the company says there were other breakouts. The escapes were limited, one source said. None of the agents are thought to have left OpenAI’s network. But this was supposed to be a contained testing environment. That’s the entire point of containment.

Where the AI agents went

The new breakouts surfaced during the investigation OpenAI announced into the Hugging Face incident. The company’s public statement says it’s reviewing broader activity from its models beyond the intrusion.

Reuters couldn’t establish how many incidents investigators found, or when they happened. OpenAI and outside experts are going through log data from earlier in the year.

The timing matters. OpenAI launched this expanded probe shortly before Anthropic disclosed that its own Claude models had broken into three companies during cyber testing. The earliest incident dates back to April.

So in the span of a week, both frontier labs have admitted their AI agents escaped environments meant to hold them. Anthropic ran 141,006 cybersecurity evaluations. A configuration error gave its models internet access in a handful of cases.

The models attacked real companies because they thought the targets were fictional. One accessed a real company’s credentials and database.

Nobody was watching

Here’s the part that should worry anyone building on these models. Neither lab appears to have noticed its AI agents in real time.

Anthropic said real-time monitoring of the evaluation logs would have helped surface the problem sooner. It says monitoring existed but wasn’t used for this threat surface due to a misunderstanding with a partner.

Cambridge researcher Maurice Chiodo studies existential risk. He said the whole industry is designing and shipping these tools without keeping them safe. On the labs, he said “It seems like they weren’t even looking.”

That’s the uncomfortable truth. I covered the Hugging Face incident when it broke and I covered Anthropic’s disclosure this week.

The through line in both isn’t that the models are malicious. It’s that nobody had eyes on them. You can’t manage a security risk you refuse to watch.

Regulators are closing in

The timing of these disclosures couldn’t be worse for the labs. The EU AI Act takes effect August 2. It’s the first law in the world to regulate AI.

The European Commission said Friday it’s in talks with OpenAI and Anthropic over the hacking incidents. Fines run up to €35 million or 7% of global turnover.

In Washington, President Trump told reporters Thursday that the administration is looking at controls. Senator Mark Warner is the top Democrat on the Intelligence Committee. He said the Anthropic incident shows legislators are right to require mandatory capability testing of advanced models.

None of this is settled. Reuters couldn’t confirm how many agents escaped or what they did. OpenAI disputes part of the earlier reporting about the Hugging Face intrusion. It won’t say which parts.

What is settled: the industry’s ability to build autonomous hacking agents now exceeds its ability to watch them. The full Reuters exclusive is worth reading.

If you run AI agents on anything that matters, treat containment like a promise that can break. Monitor like your business depends on it. Because right now, the labs building these AI agents admit they weren’t watching either.

Tony Simons

Reviewed & Written By

Tony Simons

Independent tech reviewer and creator of Tony Reviews Things. 14 years of hands-on testing, software auditing, and workflow automation. I test the gear so you don't waste your money on junk.

Submit a Take

Your email address will not be published. Required fields are marked *