Research — September 2026.

The 27 papers the radar tracked in September 2026. The current radar lives on the research page. Back to the radar →

Sep 1

Stealth-release forensics

A four-stage protocol fingerprints anonymous API models using archived launch configs, tokenizer diffs, and behavioral probes, no weight access required. Run against known releases it matched 7 of 10 exactly, and its prospective call on an unlabeled August drop landed on the right version line before the vendor confirmed it. Handy checklist if you're vetting an unlabeled endpoint before you sign a vendor contract.

Read paper · arxiv.org → Models Method Sep 1, 2026
Sep 1

Sycophancy transfers from teacher to student, and you can't filter it out

Training OLMo 3 on preference pairs from different teacher models, the authors found student sycophancy tracks teacher sycophancy almost linearly, across DPO and six other preference-optimization objectives. No single bad example is the culprit either. The signal is smeared across the whole dataset, so removing a small chunk of it doesn't help. If a fine-tune inherited an agreeable base model, this is probably why.

Read paper · arxiv.org → Models Dataset Sep 1, 2026
Sep 1

Your coding agent's token budget is measuring the wrong thing

Replaying 55 archived coding-agent trajectories, researchers found that instructions, tool outputs, and agent state compress and decay at different rates, so two agents with identical context windows can end up delivering very different context to the model. Compression tricks tuned on one task set didn't carry over to new ones.

Read paper · arxiv.org → Agents Method Sep 1, 2026
Sep 1

Exploring Collaboration between a language and a non-language agent

A benchmark of collaborative chess tasks shows that integrating continuous subagent representations directly into language models via learned state tokens outperforms text-based verbalization and scales effectively.

Read paper · arxiv.org → Evals Benchmark Sep 1, 2026
Sep 1

VibeVoice-ASR-Streaming Technical Report

A streaming, LLM-based end-to-end model unifies speaker-attributed speech recognition and diarization for low-latency real-time applications.

Read paper · arxiv.org → Multimodal Analysis Sep 1, 2026
Sep 1HF Daily Papers

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

DisCo is a research agent that distills operational knowledge into reusable skills, significantly improving autonomous ML research performance across benchmarks.

Read paper · arxiv.org → Evals Benchmark Sep 1, 2026
Sep 1HF Daily Papers

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

EarlyEval predicts agent outcomes from intermediate behavior to reduce evaluation cost by halting runs early with minimal accuracy loss.

Read paper · arxiv.org → Evals Benchmark Sep 1, 2026
Sep 1HF Daily Papers

Language Models Can Control Their Own Attention

Declarative Attention lets language models declare relevant context regions during reasoning to skip most KV cache reads, reducing attended tokens with small accuracy trade-offs.

Read paper · arxiv.org → Robotics Method Sep 1, 2026
Sep 1HF Daily Papers

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

A specialization pipeline combining curated problems, synthetic reasoning, supervised fine-tuning, and reinforcement learning trains competitive programming models that exceed top human scores on IOI benchmarks using iterative test-time…

Read paper · arxiv.org → Evals Benchmark Sep 1, 2026
Sep 1HF Daily Papers

Cliff: Learning Process Rewards from the First Mistake

Cliff improves reinforcement learning with verifiable rewards by using an off-the-shelf language model to detect the first reasoning error and shaping token-level advantages accordingly.

Read paper · arxiv.org → Evals Method Sep 1, 2026
Sep 1HF Daily Papers

PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation

PaperCompiler converts research papers into structured repository specifications that preserve method logic and cross-file consistency, improving code fidelity and reducing critical errors.

Read paper · arxiv.org → Models Method Sep 1, 2026
Sep 1HF Daily Papers

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

SolarWM provides an open framework and unified training recipe for building interactive video world models across diverse data sources and generator backbones, enabling long-horizon real-time rollouts.

Read paper · arxiv.org → Multimodal System Sep 1, 2026
Sep 1HF Daily Papers

Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration

CORD is a post-fit adapter that repairs calibrated probability vectors to exactly preserve original top-1 predictions while maintaining calibration quality.

Read paper · arxiv.org → Infra Method Sep 1, 2026
Sep 1HF Daily Papers

Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance

Stochastic gradient descent dynamics are modeled as a percolation process where architectural symmetries cause subnetworks to merge in discrete blocks, producing variance spikes resembling phase transitions, with similar trapping behavior…

Read paper · arxiv.org → Models Method Sep 1, 2026
Sep 1HF Daily Papers

Debias-SparseGPT: Bias-Aware Pruning for Large Language Models

Debias-SparseGPT reduces pruning-induced demographic bias in compressed large language models by incorporating representational debiasing during post-training sparsification.

Read paper · arxiv.org → Models Method Sep 1, 2026
Sep 1

RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning

RoboTok retrieves relevant human manipulation videos from the web using a latent motion space derived from 3D hand trajectories to improve robot policy training.

Read paper · arxiv.org → Robotics Benchmark Sep 1, 2026
Sep 1

A Common Measure of Communication for Speech Brain-Computer Interfaces

Open-vocabulary mutual information provides a unified metric to compare speech brain-computer interfaces across different vocabularies and conditions, revealing trade-offs between vocabulary coverage and decoding accuracy.

Read paper · arxiv.org → Multimodal Benchmark Sep 1, 2026
Sep 1

VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement

VeriPhy verifies generated video by compiling prompts into typed physical obligations, executing frozen expert analyses with provenance tracking, and mapping evidence to auditable three-valued verdicts.

Read paper · arxiv.org → Multimodal Benchmark Sep 1, 2026
Sep 1HF Daily Papers

Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems

The study formalizes multi-agent LLM coordination via bilevel games and stochastic memory reflection, introducing a grounded evaluation gate and SRMA algorithm with convergence guarantees, validated on SWE-bench.

Read paper · arxiv.org → Evals Benchmark Sep 1, 2026
Sep 1HF Daily Papers

ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding

ShallowStream uses shallow MLLM layers to build a lightweight streaming index, reducing latency while maintaining accurate video retrieval.

Read paper · arxiv.org → Multimodal Benchmark Sep 1, 2026
Sep 1HF Daily Papers

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

Teacher-Gated On-Policy Distillation verifies teacher reliability per prompt via verifier-scored probes, routing to dense distillation or verifier-grounded reinforcement learning to improve training efficiency and accuracy.

Read paper · arxiv.org → Infra Method Sep 1, 2026
Sep 1HF Daily Papers

Graph Machine: Towards Better Pretraining via Edges

A Graph Machine architecture uses sparse dynamic routing and differentiable pointer chasing to maintain linear state complexity, enabling efficient replacement of dense Transformer layers with minimal loss degradation.

Read paper · arxiv.org → Infra System Sep 1, 2026
Sep 1HF Daily Papers

MasterControl Seventeen Every Time

A governed analytics framework pairs language models for intent interpretation with deterministic policy execution of pre-approved programs, achieving full answer-and-evidence reliability where runtime-planning agents failed.

Read paper · arxiv.org → Robotics System Sep 1, 2026
Sep 1HF Daily Papers

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

Influence-guided response rewriting of selected training examples produces stronger and more persistent behavioral shifts in language models than conventional reweighting, highlighting the broader intervention leverage of influential data.

Read paper · arxiv.org → Infra Method Sep 1, 2026
Sep 1HF Daily Papers

Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles

Benchmarks for LLM-generated GPU kernels decide correctness with a few random inputs and a loose floating-point tolerance, and their verdicts now feed leaderboards and reinforcement-learning rewards.

Read paper · arxiv.org → Evals Benchmark Sep 1, 2026
Sep 1HF Daily Papers

ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models

Vision-Language-Action (VLA) models demonstrate strong generalization in robotic manipulation and navigation, but existing fine-tuning methods provide limited safety guarantees.

Read paper · arxiv.org → Robotics Method Sep 1, 2026