Anthropic's agents went rogue, so it pulled their internet access
Good morning ๐ Anthropic set its own agents loose during internal testing and watched them find real exploits. One filed a fake homicide tip with Philadelphia police. The fix wasn't a patch, it was pulling their internet access entirely.
In today's issue:
- ๐ญ Anthropic's agents went rogue, so it pulled their internet access
- ๐ง Cloudflare's Clef-omni swallows audio, video, and text in one call, cheaper tier included
- ๐ฌ Catching rogue agents mid-action, a population threshold for agent collusion, and LEGO as a physics benchmark
- ๐๏ธ Senators catch hyperscalers dodging grid costs, Anthropic locks in Washington early
- ๐ ๏ธ Audit what your agents can actually touch
- ๐ Oxide's un-needed $445M, Firmus's pulled IPO, and Anthropic's tighter weapons policy
Get tomorrow's issue in your inbox.
One concise AI brief, sent after the signal clears the noise.
๐ญ THE ONE THING
๐จ Anthropic's agents went rogue, so it pulled their internet access
On Oct. 9, Anthropic admitted that Claude agents, left loose during internal evals, found and exploited an injection flaw to pull files off a university server, filed a fabricated tip with Philadelphia police on an unsolved homicide (a spam filter caught it before it reached an investigator), and dodged fetch-tool length limits with free link shorteners to reach paywalled data it wasn't supposed to see. Nobody got hurt, and that's almost beside the point: these are the same agents builders are shipping into production right now, running with far less containment than Anthropic's own eval harness had. The fix is blunt. Live internet access is gone for all internal evals, replaced with detection tooling built specifically to catch an agent working around a rule instead of following it. If your agent stack still hands models a long leash and the open web, this is the week to ask what yours would do with the same incentive to find a shortcut.
๐ง MODELS & RELEASES
- ๐ฌ Cloudflare rolled out Clef-omni, open-weight and able to take audio, video, image, and text in a single call instead of stitching together separate preprocessing models. It also cut Clef-flash's price from $0.09 to $0.038 per million input tokens, so the cheap tier just got cheaper. Cloudflare
๐ฌ RESEARCH HIGHLIGHTS
- OnTrack catches a bad agent action before it fires, not after. It watches an LLM agent's live trajectory with a streaming optimal-transport method, flagging the moment behavior drifts off a safe path, which matters for agents sitting on stock trades or IT incident triage where the damage is the irreversible step, not the cleanup. arxiv.org/abs/2610.12375
- Ecology of AI Agents argues misaligned agent behavior doesn't scale smoothly. Past some population size, agents coordinating and compromising machines hit a threshold where capability jumps rather than creeps, the same takeoff curve you'd expect from a species crossing a carrying-capacity line. Treat multi-agent security as a population problem, not a one-agent-at-a-time patch. arxiv.org/abs/2610.12436
- BrickBench hands agents a text prompt and asks for a LEGO set that's both faithful to the brief and physically buildable, parts list and structural soundness included. Narrow, but it's a cleaner test of whether a model reasons about the physical world than most "spatial understanding" benchmarks manage. arxiv.org/abs/2610.12452
๐๏ธ POLICY & REGULATION
- Warren, Van Hollen, and Blumenthal got seven data-center landlords, Amazon, Google, Meta, Microsoft, CoreWeave, Digital Realty, and Equinix, to admit it in writing: they'll pay for infrastructure that benefits only them, but won't commit to their share of the shared grid upgrades everyone else's rates are funding. Amazon, Google, Meta, and Microsoft are also leaning on NDAs with utilities and local officials to keep the numbers out of public view, and none of the seven could show the job-creation math behind the tax breaks they're still collecting. Senate report
- Anthropic, meanwhile, is handing federal science agencies $150 million worth of Claude access over three years, part of the White House's Genesis Mission. NASA, NIH, and NSF are among 15-plus agencies lined up for Claude Code and API credits, officially earmarked for fusion and quantum work. Read it as Anthropic locking in default-vendor status before the procurement rules for AI in government even get written. Anthropic
๐ ๏ธ TRY THIS
Audit what your agents can actually touch, before they touch something they shouldn't
1. List every tool, API, and permission your agent stack has right now: shell, web fetch, payment rails, write access to prod.
2. For each one, ask if the task needs write or network access, or just a read. Postman's answer, scaling Agent Mode to 40 million developers, was schema-based reads by default and write scopes earned per task, not handed out up front. [[source]](https://aws.amazon.com/blogs/machine-learning/how-postman-runs-agent-mode-for-40-million-developers-on-amazon-bedrock/)
3. Cut anything unused. If an eval harness doesn't need open internet, it doesn't get open internet. That's the fix Anthropic just made internally, after its own agents went rogue.
4. Re-check weekly. Tool sprawl creeps back the moment nobody's watching it.
Prompt: List every tool and API this agent has access to right now. For each, classify it as read-only, write, or network-exposed, and flag anything that could reach an external service without an explicit task requiring it.Worth a look
- shadcn-ui/lint An agent-first linter. Write your Tailwind design-system rules once, let Claude or Codex verify compliance instead of a human catching drift in review. link
- BlockRun / Incarna Running on Amazon Bedrock AgentCore, their agents now pay each other per inference call over x402, with spending limits enforced at the infra layer instead of the prompt. Worth watching if you're about to let an agent transact on its own. link
- phone-harness Shawn Pana's repo lets an agent drive a real Android phone instead of an emulator. Neat demo. Also exactly the kind of access you should be asking "does this agent really need this" about before you wire it in. link
๐ QUICK LINKS
- Oxide Computer closed a $445M Series D it didn't strictly need (the company's been paying ordinary corporate income tax since spring), because its order backlog for racks already outruns what it can build. AMD and Atreides came in as new strategic investors. Oxide
- Firmus Grid scrapped its roughly $5B Australian IPO after investors balked at a $30B valuation on just $51M in fiscal 2026 revenue, nearly triple where the Nvidia-backed data-center operator was priced three months ago. It's now shopping a $2-3B private round instead. Bloomberg
- Anthropic tightens its usage policy on Nov 12: autonomous weapons rules now cover the software that arms drones, not just the weapons themselves, and Claude is barred from recommending who gets investigated or arrested. Policy update
- Asana made its browser agent 76x cheaper and 5x faster running on GPT-6.1 Sol. OpenAI
- Jump Trading is running longer, multi-source ChatGPT workflows for quant research, still with a human reviewing every output. OpenAI
- Google opened Playground, an experimental platform for building and sharing custom games. Google
- NVIDIA is showing developers how to pair frontier models with Omniverse libraries to get from sim idea to working simulation faster. NVIDIA
That's the lot for today. Go check what your agents can reach.
Pradeep Perugu