Mistral gives models a grep command

Mistral gives models a grep command

Good morning ๐Ÿ‘‹ An 86% jump on SEC filings didn't come from a bigger model. It came from letting the model grep the document itself.

In today's issue:

Get tomorrow's issue in your inbox.

One concise AI brief, sent after the signal clears the noise.


๐Ÿ”ญ THE ONE THING

๐Ÿ” Mistral gives models a grep command

Mistral shipped Agentic Search on Tuesday: five tools (search, open, navigate, read, grep) that let a model dig through a document the way an engineer would, instead of trusting whatever chunks a retriever handed it. On FinanceBench, 368 SEC filings, accuracy jumped from 26.7% to 86% with Mistral Medium 3.5. On OfficeQA Pro, 696 scanned Treasury bulletins, GLM-5.2 went from 6.3% to 51.9%, while p90 latency dropped as much as 39.6% and token spend fell by close to a third. It's live now in the Search Toolkit, wired into Studio and Vibe, cloud or on-prem. The jump is real and the failure mode it fixes (RAG choking on long, dense filings) is one every builder shipping document agents has hit, but these are Mistral's own benchmarks against Mistral's own baseline, so the number to watch is what happens when someone outside Mistral runs it.


๐Ÿง  MODELS & RELEASES


๐Ÿ”ฌ RESEARCH HIGHLIGHTS


๐Ÿ› ๏ธ TRY THIS

Point Mistral's Agentic Search at your own ugly documents before you trust the benchmark numbers

Today's lead is a retrieval layer, not a chatbot: instead of grabbing fixed chunks in one pass, it lets the model search, open, navigate, read, and grep its way through a document the way you would. Mistral's own claim is a jump from 26.7% to 86% correctness on FinanceBench. Worth checking against the documents you actually deal with.

1. Clone the Search Starter App and index a folder of the documents you already fight with: 10-Ks, contracts, spec sheets with buried tables.

2. Ask it five questions you'd normally answer by scrolling PDFs yourself. Let it run its own search/open/navigate/read/grep loop instead of prompting it directly.

3. Run the same five questions through whatever retrieval you use today and compare answers, not just the chunks each system pulled.

4. Check token spend on both runs. Mistral says Agentic Search cuts consumption by up to a third by stopping repeat searches once it finds the right page.

Prompt: Using the indexed document set, find the counterparty's late payment penalty in Section 4.2. Search for the filing, open it, navigate to Section 4.2, read the exact clause, and quote it before answering.

Worth a look: AWS Bedrock AgentCore's Policy Authoring takes a prose policy document and auto-formalizes it into an enforceable Dogwood policy, things like refund caps by time of day, identity-verification windows, per-account rate limits, checked against your tool schemas. It also hooks into Bedrock Guardrails, so a policy can reject a dispute filing for the simple reason that its free-text field contains a Social Security number. Once your agent is the one reading and acting on documents, this is the layer that stops it from acting on the wrong thing it found.


๐Ÿ”— QUICK LINKS


Go run the eval yourself before you trust anyone's benchmark, including this one.

Pradeep Perugu

Get inovAIte in your inbox.