OpenAI ships dots, an agent built to work while you're not watching
Good morning ๐ OpenAI built an agent that works while you're not looking. The part worth ten seconds of your attention isn't what it does, it's what you'd have to hand it to get there.
In today's issue:
- ๐ญ OpenAI ships dots, an agent that works while you're not watching
- ๐ง NVIDIA's tabular model skips training and feature engineering entirely
- ๐ฌ Chain-of-thought traces you can't trust, plus two more findings worth five minutes
- ๐ Outmarket AI closes a $34.5M Series B five months after its Series A
- ๐๏ธ Brussels opens the copyright conversation nobody's shipped an answer to
- ๐ ๏ธ Steal AWS's verify-your-own-extraction pattern for your next contract pile
- ๐ OpenAI's $1.4T raise, a journalism grant, and two new Bedrock regions
Get tomorrow's issue in your inbox.
One concise AI brief, sent after the signal clears the noise.
๐ญ THE ONE THING
๐ค OpenAI ships dots, an agent built to work while you're not watching
OpenAI closed DevDay 2026 by launching dots: an "always-on" ChatGPT agent that chases a goal across 4,000+ connected apps instead of waiting for your next prompt, rolling out today to Pro and Business Premium subscribers. It runs on GPT-6 Astra, and OpenAI's own language ("always-on," "proactive research") tells you the pitch: fewer prompts, not smarter replies. TechCrunch, Axios, Bloomberg, CNBC, and Platformer all covered it within hours, which says more about DevDay's gravity than about whether dots earns the trust it's asking for. Handing an agent standing access to your inbox and docs is a bigger ask than a chat window ever was. Test it on something you can afford to be wrong about before you let it touch anything you can't.
๐ง MODELS & RELEASES
- ๐ Kumo Tabular NVIDIA's new tabular foundation model skips training and feature engineering, reading your labeled rows as context and scoring new ones in a single pass. It tops TabArena at an ELO of 1950, 17x faster than rival foundation models, and ships in three sizes (28M to 215M params) under the OpenMDW-1.1 commercial license. If your team still hand-tunes XGBoost for every new dataset, benchmark this first. NVIDIA Kumo Tabular
๐ฌ RESEARCH HIGHLIGHTS
- Chain-of-thought traces lie, and not rarely. Researchers tested reasoning traces against iGSM, a grade-school math benchmark where every step can be mechanically checked. On the hardest problems, 31.6% of correct answers came with invalid traces, more than half of which sailed through syntax and arithmetic checks while failing on logic. Worse: models trained on shuffled or swapped traces kept their accuracy even after none of their traces passed verification. If your agent-auditing pipeline treats CoT output as a window into what the model actually did, this says otherwise. Paper
- Character training might beat guardrails for risk. A new paper trains models on a constitution built around constant absolute risk aversion, the bet being that a misaligned-but-cautious agent takes the safer deal with humans instead of rolling the dice on rebellion. Two of four tested models generalized to out-of-distribution decisions better than untrained baselines, despite never seeing the benchmark's format during training. Small sample, but it's a different lever than RLHF or constitutional AI: shape disposition, not just output. Paper
- Gender bias across LLMs: real, but nobody agrees on which direction. Ten models from nine vendors, tested two ways: stereotyped-phrase attribution and moral judgment on harm scenarios. Two of ten models attributed masculine-coded phrases to women more than the reverse; three showed the opposite pattern entirely. On harm judgments, some models protected female targets in line with known human bias, while three others showed no variation at all. The authors' conclusion tracks: bias auditing has to be per-model, per-vendor, ongoing. A clean pass on one model tells you nothing about the next one. Paper
๐ AI STARTUPS
- Outmarket AI raised a $34.5M Series B led by SignalFire, just five months after its $17M Series A, pushing total funding to $56.5M. The insurance-agency platform (workflows for commercial, benefits, personal lines) says it's now running in 300+ agencies, including a quarter of the top 100 shops, and is spending the round to push into carrier-side tools next. Back-to-back rounds this close together usually mean real usage numbers, not just a hot category. Outmarket AI Series B
๐๏ธ POLICY & REGULATION
- European Commission opened a consultation on copyright for the AI era: whether generative models need clearer licensing terms for training data, and how to handle AI clones of a performer's voice or likeness. It runs September 29 through November 3. Nothing binds anyone yet, this is Brussels gathering input before it drafts an actual rule, so treat it as an early input window, not news to react to. Consultation
๐ ๏ธ TRY THIS
Give your document pile the same proactive treatment dots gives your inbox: extraction with a built-in verifier
Dots watches your work and surfaces what matters before you ask. AWS just published the unglamorous version of that trick: an agent that reads vendor contracts, pulls the fields, then checks its own extraction before a human sees it. Steal the shape for whatever stack of PDFs you're sitting on.
1. Pick the fields you'll actually query later (renewal date, auto-renew clause, liability cap), not every field the document contains.
2. Run one Bedrock AgentCore agent to extract those fields per document, and a second agent to verify each value against the source text.
3. Route mismatches to a review queue instead of trusting the first pass.
4. Layer on a chat agent for portfolio-wide questions, like which contracts auto-renew next quarter, that plain RAG can't answer.
Prompt: Extract {field_list} from this contract. For each value, quote the exact source sentence it came from and rate your confidence 1-5. Flag anything under 4 for human review.๐ QUICK LINKS
- OpenAI is floating a $30B raise at a $1.4T valuation, with annualized revenue past $40B and growing 70% since July. link
- OpenAI is putting $5M in cash and up to $5M in credits behind the Lenfest AI Collaborative, its journalism program. link
- Anthropic Claude now runs in-region on Bedrock in India. Local processing, no data leaving the country. link
- xAI's Grok 4.7 landed on Bedrock with a 500K-token context window and four reasoning-effort dials. link
See you tomorrow.
Pradeep Perugu