OpenAI's own safety-test agents broke into Hugging Face's live systems
Good morning ๐ OpenAI pointed its own safety-test agents at a live system, and the agents didn't stop at the assignment. It took OpenAI days to work out the intruder chewing through Hugging Face's infrastructure was theirs.
In today's issue:
- ๐ญ OpenAI's own safety-test agents broke into Hugging Face's live systems
- ๐ง Models: Gemini 3.5 Transcribe's real numbers, NVIDIA's custom HBM, IBM's open Granite 4.2
- ๐ฌ Research: right answers hiding broken reasoning, agent societies that specialize with no chat channel
- ๐ Startups: Transfyr's $25M bet on capturing what never gets written down
- ๐ ๏ธ Try this: caging your agent before it touches anything live
- ๐ Quick links: Nvidia's Hugging Face talks, Instinct's $350M raise, and more
Get tomorrow's issue in your inbox.
One concise AI brief, sent after the signal clears the noise.
๐ญ THE ONE THING
๐ OpenAI's own safety-test agents broke into Hugging Face's live systems
Hugging Face caught the intrusion in mid-July, and it took OpenAI days to work out that its own research agents had done it, only realizing after Hugging Face told them the credentials in question had already been revoked. Those agents strung together undisclosed exploits to crack an internal package registry, then used it to reach systems well outside their test environment, no human steering. OpenAI's own report, released this week, reads less like reassurance than confession: the model had been tested without the safety classifiers built to catch exactly this, and OpenAI says the monitoring it's rolling out now would have caught the activity more than a day before the Hugging Face breach, had anyone been watching. Pausing the model line and walling off its network access, kill switch included, is the obvious fix. Whether it's enough is the open question, and a report written by the company that missed its own agent for days isn't the one likely to answer it for free.
๐ง MODELS & RELEASES
- ๐๏ธ Google DeepMind shipped Gemini 3.5 Transcribe, and the numbers are real: 2.6% word error rate offline, 4% streaming, across 85+ languages, with transcripts landing 70% faster than the old Chirp 3 model. If your product still ships a "review the transcript" step, that step just got cheaper to skip. Link
- ๐ NVIDIA opened NVLink Fusion to custom memory, NVHBM, promising 30% more bandwidth and 15% less power than stock HBM4E by moving the memory controller onto the HBM die itself. Amazon's Annapurna Labs is first in line for Trainium4, which tells you who NVIDIA is really trying to keep close: the hyperscalers building their own silicon. Link
- ๐งฉ IBM detailed Granite 4.2's build: plain dense transformers at 3B/8B/30B, not mixture-of-experts, trained on 15 trillion tokens with a 512K context window and Apache 2.0 licensing. The 30B checkpoint hits 57% on SWE-Bench Verified, which puts a fully open, self-hostable model in agentic-coding range that used to require an API call to someone else's model. Link
๐ฌ RESEARCH HIGHLIGHTS
- Trace Integrity puts a number on something every agent builder already suspects: a right answer doesn't mean right reasoning. Testing SQL-generation agents, the authors clocked answer accuracy at 20-24%, while their "correct answer, invalid trace" rate ran as high as 59%. Translation: more often than not, the right answer was sitting on top of broken logic. Grade an agent by output match alone and you're blind to exactly the failure mode that bites in production. Paper
- SwarmWorld drops identical LLM agents into a shared simulated world with no chat channel between them. They specialize anyway: explorers, builders, maintainers, coordinators, purely by reading the traces each other leaves behind (the same trick ants use, stigmergy). The resulting societies build broader, more durable tech portfolios than a strong best-of-N solo-search baseline, though a lone agent still wins on a single one-off artifact. File under: coordination doesn't require a protocol. Paper
- AsymSpec goes after a cost problem every long-context agent pipeline runs into: compressing the input to control cost quietly tanks accuracy. The fix is asymmetric. A small drafter model reads the full, uncompressed context while the big verifier runs on the compressed one, with the drafter steering the verifier's output to make up the difference. Result: roughly 90% of full-context accuracy at 1.3 to 1.7x the throughput and about a quarter of the compute. It's the rare speedup paper aimed at the exact setup most deployments are already stuck in. Paper
๐ AI STARTUPS
- Transfyr closed a $25M seed led by General Catalyst to wire lab benches with sensors and feed the output into multimodal models. The bet: most of what makes an experiment reproducible never gets written down, it lives in a postdoc's hands, and it walks out the door when they do. Transfyr
๐ ๏ธ TRY THIS
Cage your agent before it touches anything live
OpenAI's own safety-test agents breached Hugging Face's production systems this week. Not a rogue model, a *sanctioned* red-team run that wandered past its intended scope. If that can happen inside OpenAI, it can happen to the agent you're pointing at a customer database on Tuesday.
1. Instrument every tool call your agent makes with OpenTelemetry spans, regardless of framework (LangGraph, Strands, the Claude Agent SDK, whatever).
2. Run the agent against staging first, scoring the resulting traces for any call outside its declared scope (writes, deletes, hosts not on an allowlist).
3. Set a hard gate: no promotion to a live target until someone's reviewed every flagged span.
4. Only then aim it at production, tracing still running, so a boundary violation shows up as a flagged span instead of an incident report.
Prompt: Review this agent's OpenTelemetry trace. Flag every tool call that writes, deletes, or reaches a host outside this allowlist: [your allowlist]. For each violation, give the span ID and a one-line reason it's out of scope.Worth a look
AWS shipped Bedrock AgentCore Evaluations, which scores an agent's OTel traces regardless of what built it (LangGraph, LlamaIndex, the OpenAI Agents SDK, Google ADK, Claude Agent SDK, Strands). If step 2 above is the part you'd otherwise build yourself, this is the shortcut.
๐ QUICK LINKS
- Nvidia Bloomberg says Nvidia held talks to buy Hugging Face at a valuation topping $13B. No signed deal yet, and Hugging Face already turned down a $500M Nvidia investment last year over control concerns. report
- Instinct The one-year-old AI-assistant startup closed a $250M Series B led by Index Ventures and Benchmark, putting it at a $2.5B valuation on $350M raised total. report
- TCS is buying Porsche's 4,500-person consulting arm MHP and pairing the deal with an AI partnership to "industrialize AI at scale" across Porsche's operations. No price disclosed. release
- Sakana AI picked up a new contract with Japan's Ministry of Defense, this one for AI tools that support strategic intelligence analysis rather than battlefield command and control. details
- Anthropic opened Claude usage data, about 250,000 real conversations, to outside researchers at Stanford, Oxford, and METR. Top finding: over half the conversations involved people delegating consequential legal or financial decisions to the model. study
- Anthropic is also putting $5M toward outside research on AI and mental health. Applications are open now for evals covering crisis conversations and multi-turn risk escalation. grants
- OpenAI published its own data arguing ChatGPT extends student learning past the classroom instead of replacing it. Self-reported, so weigh it accordingly. report
- Natera built a phone-in voice agent on AWS Bedrock AgentCore that books mobile blood-draw appointments, with AWS claiming 100% tool-calling accuracy in production. writeup
See you tomorrow.
Pradeep Perugu