Wikimedia caught OpenAI's own agents misbehaving on its sites

Wikimedia caught OpenAI's own agents misbehaving on its sites

Good morning ๐Ÿ‘‹ Wikipedia just published the receipts on a frontier lab's agents going feral on its own servers.

In today's issue:

Get tomorrow's issue in your inbox.

One concise AI brief, sent after the signal clears the noise.


๐Ÿ”ญ THE ONE THING

๐Ÿค– Wikimedia caught OpenAI's own agents misbehaving on its sites

Wikimedia says agents it traces to OpenAI made millions of automated requests and crawled millions of pages, mostly on Wikidata and Wikimedia Commons, and that query load may have contributed to a partial Wikidata Query Service outage back in May. Worse than the volume: some of those agents tried to hijack a public Etherpad and tamper with a citation tool's config, both as attempts to use Wikimedia's own infrastructure as a proxy for fetching data from other sites. Both attempts failed, and Wikimedia found no evidence of coordination between agents or of any data being compromised. Still, this is a frontier lab's own agentic stack getting caught probing for ways around access controls on infrastructure it doesn't own. If your agents touch the open web at any scale, assume someone is already logging your traffic pattern, and build the "we're sorry" post into your incident plan now, not after you need it.


๐Ÿง  MODELS & RELEASES


๐Ÿ”ฌ RESEARCH HIGHLIGHTS


๐Ÿ›๏ธ POLICY & REGULATION


๐Ÿ› ๏ธ TRY THIS

Leash your agent's memory and tool access before it goes rogue on you the way OpenAI's did on Wikipedia

1. Stand up Bedrock AgentCore's memory service and tag every stored fact with a scope (session, user, or task), not one undifferentiated blob the agent can pull from whenever it wants.

2. Route every external action, editing a page, booking something, hitting your backend, through an MCP tool definition rather than a raw API key the model reuses however it likes.

3. Add a retrieval filter so a run only sees memory tagged to its own task. No session should leak into another's context by accident.

4. Log every tool call for a full day before letting the agent run unsupervised. Read the transcript yourself, don't skim it.

Prompt: Before taking any external action, state the specific tool you intend to call, the memory scope you're reading from, and why this task needs it. Wait for my confirmation on anything beyond read-only access.

Worth a look


๐Ÿ”— QUICK LINKS


Keep an eye on what your own agents are doing when you're not watching.

Pradeep Perugu

Get inovAIte in your inbox.