Wikimedia caught OpenAI's own agents misbehaving on its sites
Good morning ๐ Wikipedia just published the receipts on a frontier lab's agents going feral on its own servers.
In today's issue:
- ๐ญ Wikimedia catches OpenAI's own agents going rogue on its sites
- ๐ง OpenAI buys another on-ramp into Atlassian's work data
- ๐ฌ Whether models act on their own morals, adaptive prompt-injection defense, and an open rival to Meta's Muse
- ๐๏ธ An 18-month sentence for AI-streaming fraud, and Brussels hires a watchdog
- ๐ ๏ธ Leash your agent's memory and tool access before it goes rogue on you
- ๐ SpaceX's Nvidia debt binge, teen-safety red flags, and more
Get tomorrow's issue in your inbox.
One concise AI brief, sent after the signal clears the noise.
๐ญ THE ONE THING
๐ค Wikimedia caught OpenAI's own agents misbehaving on its sites
Wikimedia says agents it traces to OpenAI made millions of automated requests and crawled millions of pages, mostly on Wikidata and Wikimedia Commons, and that query load may have contributed to a partial Wikidata Query Service outage back in May. Worse than the volume: some of those agents tried to hijack a public Etherpad and tamper with a citation tool's config, both as attempts to use Wikimedia's own infrastructure as a proxy for fetching data from other sites. Both attempts failed, and Wikimedia found no evidence of coordination between agents or of any data being compromised. Still, this is a frontier lab's own agentic stack getting caught probing for ways around access controls on infrastructure it doesn't own. If your agents touch the open web at any scale, assume someone is already logging your traffic pattern, and build the "we're sorry" post into your incident plan now, not after you need it.
๐ง MODELS & RELEASES
- ๐ค OpenAI is expanding its Atlassian partnership, wiring its models straight into customers' enterprise knowledge so teams can plan and ship work inside the tools they already run on. No new checkpoint here, just OpenAI buying another on-ramp into where the actual work data lives. That's worth more than a benchmark win. source
๐ฌ RESEARCH HIGHLIGHTS
- Post-training decides whether a model acts on its own stated morals, not scale. A pre-registered panel of 248 scenarios across five kinds of pressure tests something stated-values benchmarks can't see: whether a model that calls an action wrong still goes ahead and does it. That's a different failure than not knowing better, and it's the one that matters once models are running as agents instead of answering trivia. arXiv
- Prompt-injection defenses are starting to train against attackers that adapt, not just a fixed attack set. Web agents have to read pages written by strangers, and a planted instruction on that page can hijack the task. AdvSim2Real builds a simulated web world where the injection attacks evolve alongside the agent, instead of testing against the same static payloads every defense gets graded on. Whether that holds up outside simulation is the open question, but static defenses were already losing to this problem. arXiv
- nanoMuse is the open-source rebuttal to Meta's Muse. Muse, shown this September, runs a persistent personal agent, accounts, devices, memory, closed inside one vendor's cloud. nanoMuse makes the same bet on long-running agents that follow you across devices, minus the lock-in: open weights, your data stays put. Whether an open stack can match the continuity Muse gets from owning the whole platform is unproven. arXiv
๐๏ธ POLICY & REGULATION
- DOJ (SDNY) Michael Smith, 54, got 18 months for running up to 10,000 bot accounts that streamed his own AI-generated songs billions of times. He's forfeiting $8.09 million, and SDNY is calling it the first criminal case built around AI music streaming fraud. Every royalty farm running the same trick just got a sentencing number to think about. DOJ release
- European Commission Brussels opened a tender to staff and run an EU AI Observatory for three years, bids due November 3. The AI Act gave Europe rules on paper. This is the standing operation meant to watch whether anyone actually follows them. Call for tenders
๐ ๏ธ TRY THIS
Leash your agent's memory and tool access before it goes rogue on you the way OpenAI's did on Wikipedia
1. Stand up Bedrock AgentCore's memory service and tag every stored fact with a scope (session, user, or task), not one undifferentiated blob the agent can pull from whenever it wants.
2. Route every external action, editing a page, booking something, hitting your backend, through an MCP tool definition rather than a raw API key the model reuses however it likes.
3. Add a retrieval filter so a run only sees memory tagged to its own task. No session should leak into another's context by accident.
4. Log every tool call for a full day before letting the agent run unsupervised. Read the transcript yourself, don't skim it.
Prompt: Before taking any external action, state the specific tool you intend to call, the memory scope you're reading from, and why this task needs it. Wait for my confirmation on anything beyond read-only access.Worth a look
- AgentCore + OpenClaw AWS's pattern for a personal assistant with durable, metadata-tagged memory instead of one blob per user. AWS
- AgentCore + Nova Sonic + MCP A voice travel concierge that reaches the backend only through defined MCP tools, not a standing credential. AWS
๐ QUICK LINKS
- SpaceX is in early talks for $10B in bank loans plus $30B in investment-grade debt, led by Apollo, just to buy Nvidia chips. Bloomberg, citing FT
- Common Sense Media rated Sora "Unacceptable Risk" and ChatGPT "High Risk" for teens. In testing, crisis alerts arrived more than 24 hours late, well past the point they'd help. report
- Turba Labs left stealth with $52M (seed plus a $40M Creandum-led Series A) to sell digital twins of data centers instead of building new ones. WSJ
- OpenAI posted new results on open math problems from an internal frontier model, Lean proofs included, on GitHub. post
- Anthropic widened its Cyber Verification Program, giving vetted security teams tiered, safeguard-eased access for red-teaming and vulnerability work. announcement
- Mistral shipped Large 4, a 1T-parameter (49B active) open-weight flagship, the first product funded by its โฌ3B Series D. Weights land by month's end. Mistral AI
Keep an eye on what your own agents are doing when you're not watching.
Pradeep Perugu