Argo-Bench just showed the best data agent on the market clears a third of realistic tasks
Good morning ๐ The best model on the leaderboard still clears barely a third of realistic warehouse tasks. That's the good one.
In today's issue:
- ๐ญ Argo-Bench: the best data agent on the market still misses two of every three realistic tasks
- ๐ฌ A missing reasoning primitive in math models, and a 174x cut in fine-tuning memory
- ๐๏ธ California subpoenas OpenAI over cybersecurity incidents
- ๐ ๏ธ Put a verifier between your agent's claims and anyone who trusts them
- ๐ Instinct's $1B month, Mandiant's founder goes agentic, and three enterprise builds
Get tomorrow's issue in your inbox.
One concise AI brief, sent after the signal clears the noise.
๐ญ THE ONE THING
๐ฆ Argo-Bench just showed the best data agent on the market clears a third of realistic tasks
Argo-Bench drops data agents into a simulated NYC delivery warehouse: 235 tables, 7.5 billion rows, 210 tasks that end in real actions like banning an account or allocating a budget, not a SQL-grading rubric. The best of 14 frontier and open-weight models broke 95 points on just 34.8% of those tasks, averaging 59.5 overall. The sharper finding is buried in the setup note: the text-to-SQL benchmarks everyone cites to claim progress have answer keys that are frequently wrong. If your eval still scores SQL syntax instead of the action it produced, you're measuring whether the model can guess the grader's mistake, not whether it did the job.
๐ฌ RESEARCH HIGHLIGHTS
- The Missing Primitive splits math reasoning into four testable pieces (finding the approach, generating it, parsing it, executing it) and finds LLMs that ace frontier problems still stumble on the first one. Discovery, not execution, is the bottleneck. Target that specific gap with self-distillation and scores move across model sizes and benchmarks. Grading only the final answer hides all of this. arXiv
- TACO is a ternary optimizer that cuts fine-tuning memory 174x versus AdamW8bit, 27.7 GB down to 0.16 GB, by sparsifying gradient signs column-wise instead of keeping dense optimizer state. Peak training memory on a 13B model drops from 80.6 GB to 27.5 GB. Full-parameter fine-tuning of a 30B+ model now fits on a single 80GB H100, at comparable accuracy and runtime to existing methods. arXiv
๐๏ธ POLICY & REGULATION
- California's DOJ subpoenaed OpenAI on October 1 over cybersecurity incidents tied to its models, including last month's Hugging Face breach. AG Bonta isn't being subtle: labs "have a moral and legal responsibility" to stop their tools from enabling cyberattacks, and he says they "can and should be held legally accountable." Same office is already probing xAI's Grok over nonconsensual content and gearing up to enforce the state's new chatbot safety laws, SB 1119 and SB 867. Builder read: California is where the enforcement teeth are actually landing, not Washington. OAG press release
๐ ๏ธ TRY THIS
Stop trusting your agent's "done." Put a verifier between its output and anyone who acts on it.
1. Pick one agent task you don't fully trust yet (a migration script, a cleanup job, a form-filler) and list the state it's supposed to change.
2. After each run, skip the agent's summary. Query the actual state directly and diff it against what the agent claimed happened.
3. Where you can, bound the agent to a fixed set of typed operations instead of freeform shell or API calls. Smaller surface, easier to audit.
4. Log the mismatch rate for a week. Anywhere near the ~35% success rate Argo-Bench found on realistic tasks, and that agent isn't ready to run unsupervised.
Prompt: After completing the task, list every piece of state you changed (file, record, field) with before/after values. Do not summarize success or failure, only report the raw diffs.Worth a look
- AWS's Adjudicated Query pattern bounds a chat agent to six typed operations over a deterministic rules engine, so compliance answers ship with a "completeness receipt": compliant plus in-breach plus ambiguous plus unreadable has to equal the number scanned. Same instinct as today's lead: don't let the model grade its own homework. AWS ML blog
- ServiceNow's AutoSynthData hunts for an agent's actual failure modes and generates synthetic tasks to train against them. On ServiceNow's own ITSM benchmark that took Pass@1 from 18.77% to 27.18%. A way to go looking for the state-corrupting cases before a benchmark finds them for you. Hugging Face
๐ QUICK LINKS
- Instinct The consumer agent startup raised a $1B Series C at a $10B valuation, a month after its $250M Series B priced it at $2.5B. TechCrunch
- Armadin Mandiant founder Kevin Mandia's agent-swarm pentesting startup closed a $255.5M Series B at $2.5B+, with a16z and Accel co-leading. TechCrunch
- OpenAI / Chatham Financial The derivatives firm cut trade validation from 30 minutes to under 4 by building on Codex. OpenAI
- OpenAI / Albertsons The grocery chain is pushing ChatGPT Enterprise and the OpenAI API into store operations to speed up its teams. OpenAI
Check the state, not the summary, before you let an agent run unsupervised.
Pradeep Perugu