Mistral gives models a grep command
Good morning ๐ An 86% jump on SEC filings didn't come from a bigger model. It came from letting the model grep the document itself.
In today's issue:
- ๐ญ Mistral gives models a grep command
- ๐ง Liquid AI's free 3x inference speedup
- ๐ฌ Reasoning models that know when to stop thinking
- ๐ ๏ธ Try Mistral's retrieval claims on your own documents
- ๐ Broadcom's $60B chip bet, Anthropic's retention U-turn, and more
Get tomorrow's issue in your inbox.
One concise AI brief, sent after the signal clears the noise.
๐ญ THE ONE THING
๐ Mistral gives models a grep command
Mistral shipped Agentic Search on Tuesday: five tools (search, open, navigate, read, grep) that let a model dig through a document the way an engineer would, instead of trusting whatever chunks a retriever handed it. On FinanceBench, 368 SEC filings, accuracy jumped from 26.7% to 86% with Mistral Medium 3.5. On OfficeQA Pro, 696 scanned Treasury bulletins, GLM-5.2 went from 6.3% to 51.9%, while p90 latency dropped as much as 39.6% and token spend fell by close to a third. It's live now in the Search Toolkit, wired into Studio and Vibe, cloud or on-prem. The jump is real and the failure mode it fixes (RAG choking on long, dense filings) is one every builder shipping document agents has hit, but these are Mistral's own benchmarks against Mistral's own baseline, so the number to watch is what happens when someone outside Mistral runs it.
๐ง MODELS & RELEASES
- โก Liquid AI put out DSpark, ~300M-parameter draft models that speculative-decode the LFM2.5 family. Up to 3.18x faster on an H100, 2.87x on a MacBook CPU, and the output is bit-identical to plain greedy decoding, so it's free throughput if you're already serving LFM2.5. Liquid AI
๐ฌ RESEARCH HIGHLIGHTS
- Reasoning models that size their own thinking. A 1.5B model trained with GRPO learns to pick NoThink, Short, or Long reasoning before it answers, instead of running a fixed token budget on every problem. It holds MATH500 accuracy while cutting average response length by about 41%, and the savings carry over to GSM8K on the easy questions. If you're paying per token for reasoning models, this is the fix nobody's shipped yet. arXiv
- Most "self-improvement" claims don't survive a real null test. An audit of Qwen3-8B under LoRA self-training found seven ways researchers fool themselves, the worst being a single greedy-decode ledger that manufactures capability gains on an untrained model purely from inference-batching noise. Rerun with a proper per-problem test against a measured baseline, external distillation genuinely helps on problems the base model rarely solves. Self-training alone doesn't, and it quietly breaks problems the model already had right. Ask any "self-improving model" claim what null it was tested against. arXiv
- Semantic cache eviction isn't where the money is. A clean head-to-head of seven eviction policies (FIFO, LRU, LFU, ARC, GDSF, streaming SISO, and a redundancy-aware method) across workloads and encoders found LFU edges out the rest as the default, barely. The number that matters more: at normal similarity thresholds, only 2.1 to 3.9% of LMSYS and QQP "hits" are actually answer-substitutable, which drags claimed hit rates of 51 to 60% down to 1.1 to 2.2% in practice. Verify the cached answer is right before you bother tuning the eviction policy. arXiv
๐ ๏ธ TRY THIS
Point Mistral's Agentic Search at your own ugly documents before you trust the benchmark numbers
Today's lead is a retrieval layer, not a chatbot: instead of grabbing fixed chunks in one pass, it lets the model search, open, navigate, read, and grep its way through a document the way you would. Mistral's own claim is a jump from 26.7% to 86% correctness on FinanceBench. Worth checking against the documents you actually deal with.
1. Clone the Search Starter App and index a folder of the documents you already fight with: 10-Ks, contracts, spec sheets with buried tables.
2. Ask it five questions you'd normally answer by scrolling PDFs yourself. Let it run its own search/open/navigate/read/grep loop instead of prompting it directly.
3. Run the same five questions through whatever retrieval you use today and compare answers, not just the chunks each system pulled.
4. Check token spend on both runs. Mistral says Agentic Search cuts consumption by up to a third by stopping repeat searches once it finds the right page.
Prompt: Using the indexed document set, find the counterparty's late payment penalty in Section 4.2. Search for the filing, open it, navigate to Section 4.2, read the exact clause, and quote it before answering.Worth a look: AWS Bedrock AgentCore's Policy Authoring takes a prose policy document and auto-formalizes it into an enforceable Dogwood policy, things like refund caps by time of day, identity-verification windows, per-account rate limits, checked against your tool schemas. It also hooks into Bedrock Guardrails, so a policy can reject a dispute filing for the simple reason that its free-text field contains a Social Security number. Once your agent is the one reading and acting on documents, this is the layer that stops it from acting on the wrong thing it found.
๐ QUICK LINKS
- Broadcom is chasing $60 billion-plus in new debt, backed by Blackstone and Apollo, to keep custom chips flowing to Anthropic and OpenAI. Bloomberg
- Anthropic will let enterprises hold their mandatory 30-day data retention on their own cloud instead of Anthropic's servers. Salesforce and other regulated customers pushed for it. Bloomberg
- OpenAI will now pace frontier model releases against cyber-capability risk, not just raw benchmarks. OpenAI
- OpenAI opened a new blog series, AI Futures, on how transformative AI reshapes power and governance. Agenda-setting, not a product. OpenAI
- AWS now routes GPT-5.6 (Sol, Terra, Luna) across 25+ regions on Bedrock. OpenAI's models, someone else's infrastructure. AWS
Go run the eval yourself before you trust anyone's benchmark, including this one.
Pradeep Perugu