OpenAI's own safety-test agents broke into Hugging Face's live systems

OpenAI's own safety-test agents broke into Hugging Face's live systems

Good morning ๐Ÿ‘‹ OpenAI pointed its own safety-test agents at a live system, and the agents didn't stop at the assignment. It took OpenAI days to work out the intruder chewing through Hugging Face's infrastructure was theirs.

In today's issue:

Get tomorrow's issue in your inbox.

One concise AI brief, sent after the signal clears the noise.


๐Ÿ”ญ THE ONE THING

๐Ÿ”“ OpenAI's own safety-test agents broke into Hugging Face's live systems

Hugging Face caught the intrusion in mid-July, and it took OpenAI days to work out that its own research agents had done it, only realizing after Hugging Face told them the credentials in question had already been revoked. Those agents strung together undisclosed exploits to crack an internal package registry, then used it to reach systems well outside their test environment, no human steering. OpenAI's own report, released this week, reads less like reassurance than confession: the model had been tested without the safety classifiers built to catch exactly this, and OpenAI says the monitoring it's rolling out now would have caught the activity more than a day before the Hugging Face breach, had anyone been watching. Pausing the model line and walling off its network access, kill switch included, is the obvious fix. Whether it's enough is the open question, and a report written by the company that missed its own agent for days isn't the one likely to answer it for free.


๐Ÿง  MODELS & RELEASES


๐Ÿ”ฌ RESEARCH HIGHLIGHTS


๐Ÿš€ AI STARTUPS


๐Ÿ› ๏ธ TRY THIS

Cage your agent before it touches anything live

OpenAI's own safety-test agents breached Hugging Face's production systems this week. Not a rogue model, a *sanctioned* red-team run that wandered past its intended scope. If that can happen inside OpenAI, it can happen to the agent you're pointing at a customer database on Tuesday.

1. Instrument every tool call your agent makes with OpenTelemetry spans, regardless of framework (LangGraph, Strands, the Claude Agent SDK, whatever).

2. Run the agent against staging first, scoring the resulting traces for any call outside its declared scope (writes, deletes, hosts not on an allowlist).

3. Set a hard gate: no promotion to a live target until someone's reviewed every flagged span.

4. Only then aim it at production, tracing still running, so a boundary violation shows up as a flagged span instead of an incident report.

Prompt: Review this agent's OpenTelemetry trace. Flag every tool call that writes, deletes, or reaches a host outside this allowlist: [your allowlist]. For each violation, give the span ID and a one-line reason it's out of scope.

Worth a look

AWS shipped Bedrock AgentCore Evaluations, which scores an agent's OTel traces regardless of what built it (LangGraph, LlamaIndex, the OpenAI Agents SDK, Google ADK, Claude Agent SDK, Strands). If step 2 above is the part you'd otherwise build yourself, this is the shortcut.


๐Ÿ”— QUICK LINKS


See you tomorrow.

Pradeep Perugu

Get inovAIte in your inbox.