Anthropic's bio-weapons filter was dark for 11 months. They published the receipts anyway.

Anthropic's bio-weapons filter was dark for 11 months. They published the receipts anyway.

Good morning ๐Ÿ‘‹ Anthropic just admitted its bio-weapons classifier sat dark for eleven months while 133 million conversations passed through unchecked. The lab published the number anyway, and the aftermath is stranger than the outage.

In today's issue:

Get tomorrow's issue in your inbox.

One concise AI brief, sent after the signal clears the noise.


๐Ÿ”ญ THE ONE THING

๐Ÿงฌ Anthropic's bio-weapons filter was dark for 11 months. They published the receipts anyway.

For nearly a year, from May 2025 to this April, the classifiers meant to stop Claude from helping build a biological weapon never touched the 133 million exchanges running through Anthropic's human-feedback platform, some 50,000 contractors' worth of conversation flying past a filter that wasn't there. Not a failure exactly, a gap: flagged traffic, per the report, "was not recorded or propagated to any review mechanisms." A retrospective Sonnet 5 sweep of that window flagged 1,197 transcripts as high-risk, and the follow-up undercuts the scare: 757 were Anthropic's own staff testing the system, the rest mostly red-team exercises, and the 62 genuine outside cases they hand-checked showed nothing concerning. The same report quietly shelves an internal model, "Model 2," for not clearing the full predeployment suite, and nudges the misalignment-risk estimate from "very low" to "low," a change Anthropic frames as "uncertainty, not new evidence" rather than a new finding. I'll take that framing at something close to face value: a lab publishing its own near-miss, with the real numbers attached, tells you more than a competitor's silence does.


๐Ÿ”ฌ RESEARCH HIGHLIGHTS


๐Ÿš€ AI STARTUPS


๐Ÿ› ๏ธ TRY THIS

Pull a sample of what your classifier waved through, and check whether it was actually right

The lead today is a reminder that a classifier's failure mode isn't the flag it raises. It's the flag it never raises. A bad model doesn't show up as an error spike, it shows up as a clean-looking pass rate, for eleven months, across 133 million exchanges, until someone finally checks.

1. Pull a random, stratified sample of records your classifier marked "pass" over the last quarter, weighted toward your highest-volume categories.

2. Re-score that sample with a different model, ideally a different vendor, and diff the two label sets.

3. Route every disagreement to a human, and log which model was wrong and by how much.

4. If the miss rate isn't zero, make this a recurring job. Weekly for a high-volume classifier, monthly for everything else.

Prompt: Given this record and the label our production classifier assigned, independently classify it from scratch. State your label, your confidence, and one sentence on what evidence would flip your answer. Do not defer to the existing label.

Worth a look: dbx is a 20MB client for 70+ databases (Postgres, MongoDB, Redis, DuckDB, SQL Server, and more) with a built-in AI assistant and an MCP server. If the records you need for step 1 are scattered across a few different stores, it's a fast way to pull the sample without standing up a separate connection for each.


๐Ÿ”— QUICK LINKS


Check your own classifiers before someone else does.

Pradeep Perugu

Get inovAIte in your inbox.