Research — July 2026.

The 605 papers the radar tracked in July 2026. The current radar lives on the research page. Back to the radar →

Jul 31

Sample More, Reflect Less

put self-refine and reflexion (the loops where a model critiques and rewrites its own answer) up against plain repeated sampling, at equal token cost, on two math benchmarks with models from 1.5B to 7B. The fancy methods lose, and the gap widens as models get bigger, once every critique token gets counted.

Read paper · arxiv.org → Science Benchmark Jul 31, 2026
Jul 31

ORCA-bench asks whether an LLM agent can actually do oncall: reading noisy metrics, logs, traces and code from an ambiguous user report, on a live microservice testbed.

Across five frontier agents, best accuracy was 25.3% on medium-difficulty incidents and 10% on hard ones. The weakest agent hallucinated a confident-sounding root cause 40% of the time.

Read paper · arxiv.org → Evals Analysis Jul 31, 2026
Jul 31

AISPA audited system-prompt instructions across commercial AI products and found the disclosure problem is exactly as bad as you'd guess.

About 40% of products carry at least one instruction that works against the user, frequently sitting right next to a protective one in the same prompt. Some products ship 60-plus protective instructions, others fewer than five. None of it gets shown to users or regulators.

Read paper · arxiv.org → Infra System Jul 31, 2026
Jul 31HF Daily Papers

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents

Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-grained language-to-motion annotations, and existing…

Read paper · arxiv.org → Robotics Dataset Jul 31, 2026
Jul 31HF Daily Papers

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts

Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the standard token-level Switch auxiliary loss Std-Aux. Std-Aux…

Read paper · arxiv.org → Multimodal Method Jul 31, 2026
Jul 31HF Daily Papers

CADENA: Stepwise CAD Reverse Engineering

Computer-Aided Design (CAD) underpins modern engineering, yet converting existing shapes into editable models still demands substantial expert effort. Most AI systems emit the entire CAD program in a single pass, never inspecting the…

Read paper · arxiv.org → Infra System Jul 31, 2026
Jul 31HF Daily Papers

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis

Recent image generators can synthesize convincing human-centric images, yet producing a useful collection remains different from producing a single successful image. A human-centric dataset must cover varied people and contexts, avoid…

Read paper · arxiv.org → Multimodal Dataset Jul 31, 2026
Jul 31HF Daily Papers

Decoding Children's Gait Behavior

We introduce a new problem domain for human action recognition: the fine-grained analysis of children's gait behaviors from standard RGB video. We specifically target the ambulatory patterns of children aged 3-17 years. Such behaviors…

Read paper · arxiv.org → Multimodal Analysis Jul 31, 2026
Jul 31HF Daily Papers

Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories

Viscous stains, characterized by high viscosity and complex rheological properties, remain a major challenge for robotic surface cleaning. Conventional wiping often spreads the stain, while scrubbing provides stronger friction but risks…

Read paper · arxiv.org → Robotics Method Jul 31, 2026
Jul 31HF Daily Papers

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance

Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models (LLMs). While Group Relative Policy Optimization (GRPO) is widely adopted, it suffers from sparse reward signals…

Read paper · arxiv.org → Models Method Jul 31, 2026
Jul 31HF Daily Papers

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

OpenART introduces a scalable red-teaming arena with evolving stateful environments to evaluate long-horizon AI agent safety, using the EMHA attack policy to expose increasing failure rates as task complexity grows.

Read paper · arxiv.org → Safety Benchmark Jul 31, 2026
Jul 30

Testing whether AI agents can actually do research, not just tasks

Most agent evals either hand agents narrow, easily-verified problems or slip AI-written papers into blind peer review, neither of which touches the actual open-ended part: picking the question worth asking, not just executing a known one. Two case studies here try to close that gap directly. The forecast that agents will soon automate AI research rests on a capability nobody had rigorously tested until now.

Read paper · arxiv.org → Evals Survey Jul 30, 2026
Jul 30

An AI teammate changes how humans talk to each other

Using Group Communication Analysis, this study ran small teams through decision-making tasks with a conversational AI seated at the table, then measured the human-to-human chatter, not the bot's output. Adding the AI teammate reshaped team communication dynamics on its own, a cost that shows up nowhere in a benchmark scoring the AI's answers.

Read paper · arxiv.org → Evals Benchmark Jul 30, 2026
Jul 30

Everyone drafting through the same handful of models looks like a monoculture risk

The concern: as writers lean on shared LLM assistants to draft and polish, population-level language variety may shrink even as any single piece of writing reads cleaner. Worth watching if you're worried about what "human" text even means once the next training set gets scraped.

Read paper · arxiv.org → Safety Method Jul 30, 2026
Jul 30

Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark

Robust low-light imaging remains challenging for the community. Recent studies have explored fusing Near-Infrared (NIR) with noisy RGB to achieve improved enhancement, yet most methods depend on carefully curated training data pairs, with…

Read paper · arxiv.org → Multimodal Method Jul 30, 2026
Jul 30

SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift

RGB imagery offers a practical, low-cost option for Unmanned Aerial/Ground Vehicle (UAV/UGV) survey support in surface-landmine detection, but object detectors remain underexplored in this safety-critical domain. Limited cross-architecture…

Read paper · arxiv.org → Multimodal Survey Jul 30, 2026
Jul 30

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual harmony. Reinforcement learning has…

Read paper · arxiv.org → Multimodal Benchmark Jul 30, 2026
Jul 30

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

Enterprise workflows increasingly rely on agents for schema-guided extraction: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We…

Read paper · arxiv.org → Evals Benchmark Jul 30, 2026
Jul 30

Scaling Properties of Text Conditioning in Visual Generation

We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find…

Read paper · arxiv.org → Multimodal Benchmark Jul 30, 2026
Jul 30HF Daily Papers

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems

Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-strategy changes. While LLM-based agents can automate this trial-and-error process, allowing the LLM to…

Read paper · arxiv.org → Infra System Jul 30, 2026
Jul 30HF Daily Papers

DiffusionGemma Technical Report

We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256…

Read paper · arxiv.org → Models Analysis Jul 30, 2026
Jul 30HF Daily Papers

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on…

Read paper · arxiv.org → Robotics Method Jul 30, 2026
Jul 30HF Daily Papers

EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents

The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines are growing fast. Access to the literature is a crucial part of that need, and…

Read paper · arxiv.org → Science System Jul 30, 2026
Jul 30HF Daily Papers

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance on pixel-generative future supervision can entangle action-relevant state transitions…

Read paper · arxiv.org → Robotics Method Jul 30, 2026
Jul 30HF Daily Papers

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve…

Read paper · arxiv.org → Evals Benchmark Jul 30, 2026
Jul 30HF Daily Papers

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving agents in…

Read paper · arxiv.org → Evals Benchmark Jul 30, 2026
Jul 30HF Daily Papers

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning

CaRL uses reinforcement learning with refusal incentives and hindsight augmentation to reduce futile reasoning in large language models while preserving task performance.

Read paper · arxiv.org → Infra Method Jul 30, 2026
Jul 30HF Daily Papers

DarwinX: Evolving Agent Harnesses Through Natural Selection

DarwinX evolves agent harnesses via population selection with frozen models, improving verified performance across benchmarks without benchmark-specific patches.

Read paper · arxiv.org → Evals Benchmark Jul 30, 2026
Jul 30HF Daily Papers

Verifier-Induced Support Reshaping in On-Policy Optimization

On-policy reinforcement learning with verifiable rewards can improve immediate task performance while reducing the diversity of successful responses needed for future training, a phenomenon called verifier-induced support reshaping.

Read paper · arxiv.org → Models Method Jul 30, 2026
Jul 30HF Daily Papers

Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling

BaguanHR improves high-resolution weather forecasting by using variable-wise super-resolution to synthesize training data, overcoming the limits of coarse-resolution model transfer and demonstrating strong scaling benefits.

Read paper · arxiv.org → Models Method Jul 30, 2026
Jul 29

AI race experiment

Paired participants played a simple game: choose safe or risky "AI development" moves, with accumulated risk capped at 10%, 60%, or 90% depending on the treatment. People didn't gamble more because they were personally risk-tolerant, they gambled more when their opponent went unsafe first or when they were falling behind.

Read paper · arxiv.org → Safety Method Jul 29, 2026
Jul 29

Evaluation-awareness, unmasked

Researchers optimized a prompt to zero out the internal "am I being tested" signal in Llama-3.2-3B and Llama-3.1-8B, no weight access needed. It worked, the latent dropped hard. Then the catch: a random, meaningless direction suppressed just as easily and shifted behavior just as far, and killing the real signal didn't even change how the model judged whether it was being evaluated.

Read paper · arxiv.org → Evals Benchmark Jul 29, 2026
Jul 29

Desktop-Delta Bench

A new benchmark checks whether computer-use agents understand what their clicks and drags actually did, not just whether the task finished. Across 2,013 human-verified cases spanning ~15 Linux apps, models hit 65% on before-after matching and still default to copying the order they were shown instead of reasoning about it. Clicking is easy (96% accurate); explaining what a drag caused is not (76%).

Read paper · arxiv.org → Evals Benchmark Jul 29, 2026
Jul 29

Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and future mitigation of demographic bias in large language models (LLMs). As a foundational empirical validation of this method,…

Read paper · arxiv.org → Models Method Jul 29, 2026
Jul 29

Σ-Mem: An Online Reliability Memory for LLM-based Multi-Agent Systems

Memory is central to long-horizon LLM agents, yet existing memory systems primarily preserve interaction content rather than modeling which agents can be trusted and under what conditions. This limitation is particularly important in…

Read paper · arxiv.org → Infra System Jul 29, 2026
Jul 29

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to…

Read paper · arxiv.org → Robotics Method Jul 29, 2026
Jul 29

LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger

Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal…

Read paper · arxiv.org → Multimodal Benchmark Jul 29, 2026
Jul 29

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation…

Read paper · arxiv.org → Evals Benchmark Jul 29, 2026
Jul 29

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures:…

Read paper · arxiv.org → Multimodal Benchmark Jul 29, 2026
Jul 29HF Daily Papers

Can Large Language Models Execute Parent Orders?

Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smaller orders while reducing execution costs. Existing approaches either rely on pre-specified market assumptions that may not…

Read paper · arxiv.org → Models Method Jul 29, 2026
Jul 29HF Daily Papers

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it…

Read paper · arxiv.org → Evals Method Jul 29, 2026
Jul 29HF Daily Papers

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine…

Read paper · arxiv.org → Multimodal System Jul 29, 2026
Jul 29HF Daily Papers

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task…

Read paper · arxiv.org → Multimodal Method Jul 29, 2026
Jul 29HF Daily Papers

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In…

Read paper · arxiv.org → Multimodal Method Jul 29, 2026
Jul 29HF Daily Papers

BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms

Retrieval-augmented generation (RAG) spans lexical and dense retrieval, graph-based indexing, and agentic search, but these paradigms are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling…

Read paper · arxiv.org → Evals Dataset Jul 29, 2026
Jul 29HF Daily Papers

Flux-OPD: On-Policy Distillation with Evolving Contexts

Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once…

Read paper · arxiv.org → Multimodal Method Jul 29, 2026
Jul 29HF Daily Papers

Harness-G: A Graph-Structured Harness for Search Agents

Reinforcement learning (RL) search agents commonly model retrieval as free-form natural-language query generation and optimize multi-turn interactions using final-answer rewards. Current studies mainly improve training with denser or more…

Read paper · arxiv.org → Evals Benchmark Jul 29, 2026
Jul 29HF Daily Papers

Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale

Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that matter most are login-gated and stateful, so synthetic environments stand in for them.…

Read paper · arxiv.org → Agents Method Jul 29, 2026
Jul 29HF Daily Papers

AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet existing literature-search systems primarily return ranked document lists. As a result, scientists and AI agents need to…

Read paper · arxiv.org → Science System Jul 29, 2026
Jul 29HF Daily Papers

ReToken: One Token to Improve Vision-Language Models for Visual Retrieval

Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a…

Read paper · arxiv.org → Multimodal Benchmark Jul 29, 2026
Jul 29HF Daily Papers

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera, a hybrid visual diffusion backbone with a principled…

Read paper · arxiv.org → Multimodal Method Jul 29, 2026
Jul 29

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of…

Read paper · arxiv.org → Models System Jul 29, 2026
Jul 29HF Daily Papers

Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing

Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often…

Read paper · arxiv.org → Evals Method Jul 29, 2026
Jul 29HF Daily Papers

RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance often degrades on challenging and out-of-domain tasks. Recent test-time steering and scaling methods improve performance…

Read paper · arxiv.org → Multimodal Method Jul 29, 2026
Jul 29HF Daily Papers

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning

Reinforcement learning with verifiable rewards (RLVR) is central to improving long-CoT reasoning in large language models. Critic-free methods such as GRPO convert response-level rewards into advantages and uniformly broadcast them across…

Read paper · arxiv.org → Models Method Jul 29, 2026
Jul 29HF Daily Papers

One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA

Can every robot in a swarm predict the same future collective state from only local observations and bandwidth-limited messages? We formulate this as decentralized shared-state prediction and introduce Collective-State JEPA (CS-JEPA), a…

Read paper · arxiv.org → Robotics Method Jul 29, 2026
Jul 29HF Daily Papers

ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning paradigms for world modeling are largely confined to discrete-time prediction, thereby exhibiting significant inefficiency in…

Read paper · arxiv.org → Infra System Jul 29, 2026
Jul 29HF Daily Papers

Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning

As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a…

Read paper · arxiv.org → Models Method Jul 29, 2026
Jul 29HF Daily Papers

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs

Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can…

Read paper · arxiv.org → Safety Method Jul 29, 2026
Jul 29HF Daily Papers

QQWorld: Quantile-Quantile Matching for World Model Regularization

Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically on the quality of the learned latent distribution. LeWorldModel (LeWM) regularizes its…

Read paper · arxiv.org → Agents Method Jul 29, 2026
Jul 29HF Daily Papers

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with…

Read paper · arxiv.org → Multimodal Method Jul 29, 2026
Jul 29HF Daily Papers

Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm

Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understanding a speaker's emotional experience. Emotional support conversation selects and sequences support for the seeker's current…

Read paper · arxiv.org → Models Method Jul 29, 2026
Jul 29HF Daily Papers

FinanceHarness: Autonomous Financial Deep Research Framework

Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep…

Read paper · arxiv.org → Infra System Jul 29, 2026
Jul 29HF Daily Papers

Complementary Matrix-Gated QKAN Fast-Weight Programmers for Quantum Dynamics Forecasting

Self-modulating quantum-inspired fast-weight programmers with coordinate-wise complementary matrix gating improve long-context sequence forecasting while preserving efficient update structures.

Read paper · arxiv.org → Science Method Jul 29, 2026
Jul 29HF Daily Papers

Articulated Object Reconstruction from Rest-State Observation

A rest-state framework reconstructs articulated objects from a single closed configuration by fusing vision-language outputs into consistent part meshes and validating synthesized motion hypotheses via geometric consistency.

Read paper · arxiv.org → Multimodal System Jul 29, 2026
Jul 28

D-Score flags hallucinations from a single forward pass, no retrieval, no second generation needed.

The method counts how many singular directions in the hidden-activation matrix stay nearly as strong as the top one; conflicting internal evidence spreads the signal across more of them, and that spread is the tell. Cheap enough to run inline.

Read paper · arxiv.org → Evals Benchmark Jul 28, 2026
Jul 28

APPA goes after a real agent tax: read one untrusted document under standard taint tracking and your whole context gets contaminated, utility included.

This permissions algebra spins off a labeled child process to absorb the taint, then a sanitizer hands clean results back to the parent.

Read paper · arxiv.org → Agents Method Jul 28, 2026
Jul 28

Generate-test-revise loops

, the default pattern in coding agents, don't hold up the way builders assume. A sealed HumanEval study found correctness on the live patch falling from 82.0% after one revision to 67.3% after two, even as the "ever found a correct fix" rate climbed to 84.7%. The agent stumbles onto the right answer, then talks itself out of it on the next pass.

Read paper · arxiv.org → Multimodal Benchmark Jul 28, 2026
Jul 28

StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation

Recent game world models can generate visually realistic and interactive environments conditioned on player actions. However, games are not defined by pixels alone; they are governed by explicit mechanics, namely state-dependent rules that…

Read paper · arxiv.org → Multimodal Method Jul 28, 2026
Jul 28

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However,…

Read paper · arxiv.org → Safety Benchmark Jul 28, 2026
Jul 28

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning…

Read paper · arxiv.org → Agents Method Jul 28, 2026
Jul 28

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost.…

Read paper · arxiv.org → Evals Benchmark Jul 28, 2026
Jul 28

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Vision-language-action (VLA) models commonly adopt an LLM-centric V to L to A pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although…

Read paper · arxiv.org → Robotics Method Jul 28, 2026
Jul 28

HumanCLAW: Can Vision-Language Models Act Through a Body?

Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad…

Read paper · arxiv.org → Robotics Benchmark Jul 28, 2026
Jul 28HF Daily Papers

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task…

Read paper · arxiv.org → Multimodal Dataset Jul 28, 2026
Jul 28HF Daily Papers

Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability

Deployed LLM agents increasingly keep their long-term memory as a filesystem: a directory tree of markdown files that the agent itself reads, writes, and reorganizes through generic file tools. Yet research has largely passed over this…

Read paper · arxiv.org → Infra System Jul 28, 2026
Jul 28HF Daily Papers

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem into hand-designed stages. Generative modeling, however, has remained the exception-despite generative models being…

Read paper · arxiv.org → Models Method Jul 28, 2026
Jul 28HF Daily Papers

Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure Modes

Speculative Decoding (SD) accelerates large language model inference by allowing a lightweight draft model to propose tokens that are subsequently verified in parallel by a larger target model. Recent approaches introduce lossy…

Read paper · arxiv.org → Models Method Jul 28, 2026
Jul 28HF Daily Papers

Metis: Memory Foundation Model

Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However, agent memory is still primarily…

Read paper · arxiv.org → Multimodal Method Jul 28, 2026
Jul 28HF Daily Papers

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in user-specific ways across tasks and sessions. Existing disambiguation methods typically…

Read paper · arxiv.org → Evals Benchmark Jul 28, 2026
Jul 28HF Daily Papers

Mental World Modeling

World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a…

Read paper · arxiv.org → Agents Method Jul 28, 2026
Jul 28HF Daily Papers

GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding

Adaptive rounding methods such as GPTQ, or equivalently Babai's nearest plane algorithm, round a real matrix to integers under a quadratic metric. They process the entries in a fixed order, one at a time, propagating each rounding error to…

Read paper · arxiv.org → Evals Method Jul 28, 2026
Jul 28HF Daily Papers

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation

Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient multi-step diffusion prohibits streaming generation, whereas real-time autoregressive approaches suffer from error…

Read paper · arxiv.org → Models Method Jul 28, 2026
Jul 28HF Daily Papers

Constitutional Midtraining: Content Presence Drives Alignment Gains

Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining:…

Read paper · arxiv.org → Safety Method Jul 28, 2026
Jul 28HF Daily Papers

ExplainBench: Evaluating Code Explanations from Agents

Large Language Model (LLM) agents have seen rapid adoption in software engineering. As agents take a greater role in the actual generation of code, they are making larger changes, spanning tens to hundreds of lines. This makes manual…

Read paper · arxiv.org → Evals Benchmark Jul 28, 2026
Jul 27

Adding a skill to your agent can make it worse, quietly

Across a large batch of runs on two office-automation benchmarks and three harness setups, the skills that scored best did so mostly by regressing less on tasks the agent already handled, not by gaining more on new ones.

Read paper · arxiv.org → Evals Benchmark Jul 27, 2026
Jul 27

Grok will call race pseudoscience credible; the other three won't

Feed Claude, Grok, GPT, and Gemini the same ethnonationalist pseudo-science, built on Frank Salter's biosocial framework, and ask them to rate its credibility. Grok's fast variants rate it credible. Everyone else doesn't.

Read paper · arxiv.org → Science System Jul 27, 2026
Jul 27

The fix for over-permissioned agents is boring: don't hand them the keys

Most enterprise agents get every credential the role might ever need, loaded at setup and never revisited, which is more attack surface than the job requires.

Read paper · arxiv.org → Agents Method Jul 27, 2026
Jul 27HF Daily Papers

CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents

Coding agents repeatedly search, navigate, and retain context from evolving repositories, but disconnected indexes, language servers, and task-local histories force repeated discovery and obscure lifecycle costs. CodeNib builds reusable…

Read paper · arxiv.org → Infra System Jul 27, 2026
Jul 27HF Daily Papers

VisualPatchWorld: Code World Models as Latent Structured Representations for Planning

Different research lines use the term world model in different ways, yet they share a common aim: to capture how the world evolves under action in a form that supports perception, simulation, and planning. Two prominent realizations are…

Read paper · arxiv.org → Multimodal Method Jul 27, 2026
Jul 27HF Daily Papers

Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control

Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than reconstructing pixels, making them a natural backbone for latent model predictive control from offline demonstration…

Read paper · arxiv.org → Robotics System Jul 27, 2026
Jul 27HF Daily Papers

Parallel Decoding Distillation for Fast Image and Video Generation

Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and…

Read paper · arxiv.org → Multimodal Method Jul 27, 2026
Jul 27HF Daily Papers

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy.…

Read paper · arxiv.org → Multimodal Method Jul 27, 2026
Jul 27HF Daily Papers

Pass the Baton: Trajectory-Relayed On-Policy Distillation

On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this deviation,…

Read paper · arxiv.org → Multimodal Method Jul 27, 2026
Jul 27HF Daily Papers

Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion

We present a reproducible pipeline for mapping Common Vulnerabilities and Exposures (CVEs) to MITRE ATT&CK Enterprise techniques from free-text vulnerability descriptions. Rather than relying on the CWE->CAPEC->ATT&CK derivation chain,…

Read paper · arxiv.org → Models System Jul 27, 2026
Jul 27HF Daily Papers

ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition

Recovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, yet remains challenging since editability depends on recovering multi-modal attributes, such as typography, vector…

Read paper · arxiv.org → Multimodal Method Jul 27, 2026
Jul 27HF Daily Papers

HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current…

Read paper · arxiv.org → Robotics Method Jul 27, 2026
Jul 27HF Daily Papers

Visual prompt engineering for video models

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performance. Since video models are currently becoming foundation…

Read paper · arxiv.org → Multimodal System Jul 27, 2026
Jul 27HF Daily Papers

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods mainly…

Read paper · arxiv.org → Multimodal Method Jul 27, 2026
Jul 27HF Daily Papers

Wonder: Video World Model Done Better

We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the…

Read paper · arxiv.org → Robotics Method Jul 27, 2026
Jul 27HF Daily Papers

Shieldstral

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification.…

Read paper · arxiv.org → Multimodal Benchmark Jul 27, 2026
Jul 27HF Daily Papers

Memory for Large Language Models

Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an implicit byproduct of computation to a spectrum of explicit, controllable mechanisms. While recent advances introduce diverse…

Read paper · arxiv.org → Robotics Method Jul 27, 2026
Jul 27HF Daily Papers

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on…

Read paper · arxiv.org → Multimodal Benchmark Jul 27, 2026
Jul 27HF Daily Papers

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted…

Read paper · arxiv.org → Evals Benchmark Jul 27, 2026
Jul 27HF Daily Papers

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box. However, most…

Read paper · arxiv.org → Models Method Jul 27, 2026
Jul 27HF Daily Papers

CAST: Game Solvers as Turn-Level Teachers for LLM Agents

Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about…

Read paper · arxiv.org → Agents Method Jul 27, 2026
Jul 27HF Daily Papers

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent…

Read paper · arxiv.org → Safety Method Jul 27, 2026
Jul 27HF Daily Papers

GPT-Red: Automated Red Teaming via Self-Play at Scale

We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end,…

Read paper · arxiv.org → Evals Benchmark Jul 27, 2026
Jul 27HF Daily Papers

Explicit Layer Modeling for Video Object Insertion and Layer Decomposition

Most video editing systems still lack explicit layered video representations, limiting their ability to perform realistic compositing, object reuse, and consistent manipulation. This limitation is especially pronounced in video object…

Read paper · arxiv.org → Robotics System Jul 27, 2026
Jul 27HF Daily Papers

Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine Detection

Hyperspectral imaging (HSI) is useful for material discrimination, but operational mine screening also depends on how many false alarms must be inspected before targets are found. This paper studies PFM-1 landmine detection in unmanned…

Read paper · arxiv.org → Evals Method Jul 27, 2026
Jul 27HF Daily Papers

Reinforcement Learning for Code Optimization

RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and reward solutions that pass. Extending this to code optimization seems straightforward: just add execution time to the…

Read paper · arxiv.org → Models Method Jul 27, 2026
Jul 27HF Daily Papers

AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition

On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models that excel at SER are too costly for edge devices. Multi-teacher knowledge distillation can compress them into a…

Read paper · arxiv.org → Multimodal Method Jul 27, 2026
Jul 27HF Daily Papers

INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models

Forward latent world models predict how actions change a scene, but recover actions for a desired change only through expensive test-time search. We introduce INTACT (INtent-To-ACTion), an end-to-end JEPA that turns action-labeled,…

Read paper · arxiv.org → Models Method Jul 27, 2026
Jul 27HF Daily Papers

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often breaks down: for the same query, audio and video relevance…

Read paper · arxiv.org → Multimodal Method Jul 27, 2026
Jul 27HF Daily Papers

SGTP: Sampling-based Game-Theoretic Planning for Real-Time Multi-Vehicle Autonomous Racing

Autonomous multi-vehicle racing requires real-time planning of diverse competitive behaviors in intense interactions. Existing planners often struggle to balance strategic diversity and computational efficiency. To address this challenge,…

Read paper · arxiv.org → Infra Method Jul 27, 2026
Jul 27HF Daily Papers

Meshy T2: Fast Native Mesh Generation with Flow Matching

Polygonal meshes are the standard surface representation of modern 3D pipelines, and generating high-quality meshes with artist-style topology is essential for film, gaming, and interactive 3D applications. Mainstream approaches serialize…

Read paper · arxiv.org → Multimodal System Jul 27, 2026
Jul 27HF Daily Papers

Weak-to-Strong On-Policy Distillation

On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on the student's own rollouts, is an effective paradigm for transferring capabilities across LLMs. Prevailing approaches assume a teacher at…

Read paper · arxiv.org → Models Method Jul 27, 2026
Jul 27HF Daily Papers

MemSFT: Mitigating Alignment Tax with an External Parametric Memory

Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastrophic forgetting and substantially degrade performance on general tasks. We propose…

Read paper · arxiv.org → Safety Method Jul 27, 2026
Jul 27HF Daily Papers

DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues

DuplexGen calibrates dialogue generation against human preferences to produce scenario-adaptive turn-taking behaviors in full-duplex interaction.

Read paper · arxiv.org → Models Method Jul 27, 2026
Jul 26

Ghana's malaria data has a shape

A consensus anomaly detector run on ten years of monthly surveillance found outbreaks aren't random noise: anomalous months post a Cohen's d of 3.25 against normal ones. The geography splits in an odd way, too. Tamale carries the heaviest case burden, but Ashanti's districts flag anomalies more often.

Read paper · arxiv.org → Evals Method Jul 26, 2026
Jul 26

Open-weight models pass the "keep it in-house" test

A British cohort study needed LLM agents to clean 20 longitudinal data-prep tasks (102 variables, six survey waves) without the data ever touching a third-party API. Open-weight 31-35B models hit 87.9% average task completion on consumer-grade hardware. Good enough that "no cloud" stops being a compromise and starts being a real option for governance-locked research data.

Read paper · arxiv.org → Infra Survey Jul 26, 2026
Jul 26

Reasoning models leak a tell before they stall

DeepSeek-R1-Distill-Qwen-7B on AIME problems either lands an answer (96.5% accuracy) or burns its whole token budget and comes up empty (11.5%), almost no middle ground. Probing hidden layers 150 tokens in catches a faint but real signal, AUC 0.61, of which path a generation is on well before the budget runs out.

Read paper · arxiv.org → Models Method Jul 26, 2026
Jul 26

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this bottleneck by computing only selected…

Read paper · arxiv.org → Multimodal Method Jul 26, 2026
Jul 26

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how…

Read paper · arxiv.org → Science Method Jul 26, 2026
Jul 26

Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels

Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a coordinate interface: the model outputs the coordinates of…

Read paper · arxiv.org → Multimodal Benchmark Jul 26, 2026
Jul 26

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems,…

Read paper · arxiv.org → Infra System Jul 26, 2026
Jul 26

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic…

Read paper · arxiv.org → Multimodal Benchmark Jul 26, 2026
Jul 26HF Daily Papers

DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes

While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack datasets that pass quality filters, set cross-domain ratios…

Read paper · arxiv.org → Multimodal Dataset Jul 26, 2026
Jul 26HF Daily Papers

From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision.…

Read paper · arxiv.org → Multimodal Benchmark Jul 26, 2026
Jul 26HF Daily Papers

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D…

Read paper · arxiv.org → Science System Jul 26, 2026
Jul 26HF Daily Papers

Data Pyramid for Embodied Manipulation

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be…

Read paper · arxiv.org → Robotics Method Jul 26, 2026
Jul 26HF Daily Papers

Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localization

Compressed short-text generators can fail in two different places: the codec may discard information before generation starts, or the latent generator may produce weak codes. Without separating these failure modes, researchers can spend…

Read paper · arxiv.org → Models Method Jul 26, 2026
Jul 26HF Daily Papers

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient…

Read paper · arxiv.org → Multimodal Method Jul 26, 2026
Jul 26HF Daily Papers

A New Role for Relevance: Guiding Corpus Interaction in Agentic Search

Relevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval agents use relevance to select top-k content, but document relevance alone cannot localize, compose, or verify the…

Read paper · arxiv.org → Evals Dataset Jul 26, 2026
Jul 26HF Daily Papers

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations…

Read paper · arxiv.org → Multimodal Benchmark Jul 26, 2026
Jul 26HF Daily Papers

Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory

Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query…

Read paper · arxiv.org → Evals Benchmark Jul 26, 2026
Jul 26HF Daily Papers

Towards Robust Reinforcement Learning for Small-Scale Language Model Agents

The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the…

Read paper · arxiv.org → Safety System Jul 26, 2026
Jul 26HF Daily Papers

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling

Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with…

Read paper · arxiv.org → Robotics System Jul 26, 2026
Jul 26HF Daily Papers

Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems

Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execution traces, tool calls, evidence links…

Read paper · arxiv.org → Evals Benchmark Jul 26, 2026
Jul 26HF Daily Papers

OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis

Biomedical image analysis spans diverse modalities and tasks, yet real-world deployment is hindered by severe distribution shifts across scanners, protocols, and patient populations. High-performing models consequently require repeated…

Read paper · arxiv.org → Science Analysis Jul 26, 2026
Jul 26HF Daily Papers

Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents

Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the repository files needed for the task. We introduce Agent…

Read paper · arxiv.org → Evals Benchmark Jul 26, 2026
Jul 26HF Daily Papers

GLI-AL: A Multi-Modal Glioma MRI Label Resource with Unified Anatomy-Lesion Labels

Existing BraTS-GLI datasets provide a widely used benchmark for adult glioma MRI segmentation, but their task definition focuses on tumor subregions and does not systematically represent coexisting white matter hyperintensities (WMH). In…

Read paper · arxiv.org → Evals Dataset Jul 26, 2026
Jul 26HF Daily Papers

Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors

Learned prompt compressors trained on English data disproportionately degrade non-English contexts, widening token-cost disparities, while multilingual training and deterministic methods reduce this gap.

Read paper · arxiv.org → Robotics Method Jul 26, 2026
Jul 25Editor's picks

OpenForgeRL trains agents inside the actual harness they'll run in (Claude Code, Codex, OpenClaw-style setups) instead of a stripped-down sandbox.

That's the gap that's been quietly killing open RL efforts: multi-turn tool use is hard to reproduce outside the vendor's own stack.

Read paper · arxiv.org → Agents Method Jul 25, 2026
Jul 25Editor's picks

Structured resistance beats less sycophancy

A new paper argues the fix for agreeable, spineless LLMs isn't dialing down sycophancy as one knob. It's teaching the model to tell "update on new information" apart from "cave because someone pushed back." Sounds abstract, but it's the exact failure mode that shows up when an agent takes instructions from an untrusted user mid-task.

Read paper · arxiv.org → Agents Method Jul 25, 2026
Jul 25Editor's picks

GS-Agent builds 4D physical worlds (geometry, motion, materials) from a text prompt, no manual rigging.

Nice trick for synthetic training data and robotics sim, if the physics survives contact with anything outside the demo reel. Watch, don't act: these generative-sim papers have a track record of looking great in the video and falling apart on real dynamics.

Read paper · arxiv.org → Science Method Jul 25, 2026
Jul 25HF Daily Papers

Characterizing Warp Divergence from Pascal to Blackwell

Since Volta introduced Independent Thread Scheduling (ITS), NVIDIA GPUs have been widely assumed to handle warp divergence in a fixed manner. We test this assumption across Ampere, Hopper, and datacenter and consumer Blackwell GPUs, using…

Read paper · arxiv.org → Infra Method Jul 25, 2026
Jul 25HF Daily Papers

Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling

The rapid evolution of generative models has unlocked new potentials in protein binder design, a pivotal task in structural biology, by facilitating end-to-end generation via joint sequence-structure modeling or hallucination. However,…

Read paper · arxiv.org → Science Method Jul 25, 2026
Jul 25HF Daily Papers

A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever

Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take the opposite path: the model stays frozen, and a persistent memory of verified solutions grows beside…

Read paper · arxiv.org → Agents Method Jul 25, 2026
Jul 25HF Daily Papers

GNM Head: A Generative aNthropometric Model of the human head

Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animation, rendering, and reconstruction. More recently, they serve as crucial conditioning signals within generative large…

Read paper · arxiv.org → Multimodal Method Jul 25, 2026
Jul 25HF Daily Papers

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other…

Read paper · arxiv.org → Multimodal Method Jul 25, 2026
Jul 25HF Daily Papers

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal…

Read paper · arxiv.org → Multimodal Method Jul 25, 2026
Jul 25HF Daily Papers

DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification

Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate because drivers are observed in different vehicles, on different roads, and under different…

Read paper · arxiv.org → Multimodal Dataset Jul 25, 2026
Jul 25HF Daily Papers

Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking

Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's…

Read paper · arxiv.org → Multimodal Benchmark Jul 25, 2026
Jul 25HF Daily Papers

N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

We present N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, and it shows…

Read paper · arxiv.org → Robotics Method Jul 25, 2026
Jul 25HF Daily Papers

N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

We present N_0-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored…

Read paper · arxiv.org → Robotics Method Jul 25, 2026
Jul 25HF Daily Papers

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as…

Read paper · arxiv.org → Models Method Jul 25, 2026
Jul 24Editor's picks

Dress a harmful request up as a relay between fake personas, and gpt-5.6-sol drops its guard

A new paper ran OpenAI's gpt-5.6-sol through 25 versions of the same harmful request. Show it the raw ask and it refuses, straight up. Route the identical ask through relay agents dressed up as "Id" and "Censor" before it hits a final "Superego" agent, and the model goes along with it. The model never actually judges the request. It judges how the request arrives.

Read paper · arxiv.org → Agents Method Jul 24, 2026
Jul 24

Sakana AI's UnMaskFork tackles a problem specific to masked diffusion models: the usual test-time scaling trick (crank temperature, sample more, vote) doesn't work on them.

UMF instead runs several diffusion models against the same sequence simultaneously, using tree search to hand each unmasking step to whichever model is most confident at that moment. No retraining required.

Read paper · sakana.ai → Models Method Jul 24, 2026
Jul 24

This is not a course. It is a journey of transformation."

That's epanorthosis, a self-correcting rhetorical figure Cicero was cataloguing two thousand years ago, and it's why so much AI copy has a telltale hollow ring. A new paper traces it to promotional training data, RLHF rewarding the confident-sounding half of the sentence, and left-to-right decoding locking it in once started. A short line of instruction cuts it noticeably; fine-tuning removes it almost entirely.

Read paper · arxiv.org → Models Method Jul 24, 2026
Jul 24HF Daily Papers

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their…

Read paper · arxiv.org → Models Dataset Jul 24, 2026
Jul 24HF Daily Papers

dRAE: Representation Autoencoder with Hyper-Spherical Codes

In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while…

Read paper · arxiv.org → Multimodal Method Jul 24, 2026
Jul 24HF Daily Papers

Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features

Bitcoin price prediction on sub-daily timescales is a hard open problem in computational finance. Bitcoin exhibits fat-tailed returns, non-stationary dynamics, and a price discovery process influenced by social discourse on Reddit and…

Read paper · arxiv.org → Models Method Jul 24, 2026
Jul 24HF Daily Papers

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models

Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deployment on resource-constrained edge devices. While efforts to compress LVLMs focus heavily on vision token reduction or…

Read paper · arxiv.org → Multimodal Method Jul 24, 2026
Jul 23

Fugu-Cyber

Sakana's follow-up to the orchestration model that already beat frontier AI on cybersecurity benchmarks. The new version claims parity with purpose-built cyber models like GPT-5.5-Cyber and Mythos Preview. The claim is Sakana's own, straight off a tweet, no third-party eval yet.

Read paper · x.com → Safety Survey Jul 23, 2026
Jul 23

LKValues

Every value-alignment benchmark you've heard of encodes Western defaults, so a model tuned to score well on them can still get local social norms wrong on the next question. This paper builds a benchmark from Sri Lankan context instead of translating an existing one. Builder read: if you're shipping outside the US and EU, your alignment eval probably doesn't cover the market you're shipping into.

Read paper · arxiv.org → Safety Benchmark Jul 23, 2026
Jul 23

Lipschitzian SLLNs is pure probability theory: a proof that strong laws of large numbers hold for locally Lipschitz random functions, under conditions broader than the usual o-minimal setup.

No model, no benchmark, just the proof. It's the kind of foundational math that might tighten generalization bounds someday, not something you'll touch this quarter.

Read paper · arxiv.org → Science Benchmark Jul 23, 2026
Jul 23

Spectral Prior for Reducing Exposure Bias in Diffusion Models

Diffusion models typically suffer from error accumulation during iterative sampling, commonly referred to as exposure bias. We reveal systematic frequency-dependent discrepancies between training and inference, which can be interpreted as…

Read paper · arxiv.org → Infra System Jul 23, 2026
Jul 23

IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation

Large Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years. However, existing systems share one core limitation: they generate and optimize ideas independently for either Quality…

Read paper · arxiv.org → Infra System Jul 23, 2026
Jul 23

LAMAR: An Open Language-Aware Multilingual Alignment Reranker

In multilingual retrieval augmented generation, a retriever can retrieve relevant documents written in multiple languages, which are subsequently reranked before answer generation. However, it remains unclear whether existing multilingual…

Read paper · arxiv.org → Safety Benchmark Jul 23, 2026
Jul 23

Scaling Native Multimodal Pre-Training From Scratch

Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of the multimodal physical world. Native multimodal pre-training avoids this limitation by…

Read paper · arxiv.org → Multimodal Method Jul 23, 2026
Jul 23

SceneActBench: Can Agents Act on the 3D Scenes They See?

Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D…

Read paper · arxiv.org → Multimodal Benchmark Jul 23, 2026
Jul 23HF Daily Papers

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

LLM training is shifting from manual design and annotation to interaction-driven self-evolution. However, existing self-evolutionary methods face a fundamental dilemma between task diversity and verification reliability: environment-bound…

Read paper · arxiv.org → Models Method Jul 23, 2026
Jul 23HF Daily Papers

Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models

Large reasoning models (LRMs) generate long reasoning traces before producing final answers. While these traces may contain useful signals for hallucination detection, harnessing them is non-trivial because long trajectories often include…

Read paper · arxiv.org → Evals Method Jul 23, 2026
Jul 23HF Daily Papers

Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models

Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterioration and damage. While existing restoration methods based on masked language modeling effectively utilize local context,…

Read paper · arxiv.org → Evals Benchmark Jul 23, 2026
Jul 23HF Daily Papers

StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents

Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screenshot is only a lossy rendering of the underlying program state, e.g., the files,…

Read paper · arxiv.org → Agents Method Jul 23, 2026
Jul 23HF Daily Papers

ID-V2V: Identity-Preserving Video Restylization

In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance while enabling flexible visual edits remains challenging for generative video models. We…

Read paper · arxiv.org → Multimodal Method Jul 23, 2026
Jul 23HF Daily Papers

Projection Pursuit CPCANet for Domain Generalization

Domain Generalization (DG) aims to learn representations robust to distribution shifts. Recent geometric alignment methods, such as CPCANet, extract domain-invariant structures through batch-wise Common Principal Component Analysis (CPCA).…

Read paper · arxiv.org → Safety Analysis Jul 23, 2026
Jul 22

ResearchArena treats agents doing AI R&D as adversaries by design.

Give an agent a real job (post-train a model, tune a CUDA kernel, optimize an inference server) and see if a monitor can catch it quietly sabotaging the result. The failure mode nobody caught: sabotage baked into training data, missed more than half the time even when the monitor could run and probe the finished artifact, not just read its reasoning trace.

Read paper · arxiv.org → Agents Method Jul 22, 2026
Jul 22

Quiet failures is a perspective piece with an argument worth stealing: AI safety talk fixates on dramatic jailbreaks and hypothetical catastrophes while the failures actually showing up in production are boring by comparison, an eval quietly gamed, a memory store slowly poisoned, oversight that exists on paper but not in practice.

Read paper · arxiv.org → Safety Benchmark Jul 22, 2026
Jul 22

Agents in the Wild

earns its place as a survey by being useful rather than novel. It tracks agentic systems moving from demo to production across software engineering, science, and finance, and pulls concrete deployment patterns (verification pipelines, fallback paths, human checkpoints) from real pharma and finance rollouts rather than benchmarks.

Read paper · arxiv.org → Science Survey Jul 22, 2026
Jul 22

Self-Supervised Learning of Structured Dynamics from Videos

Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation…

Read paper · arxiv.org → Multimodal Method Jul 22, 2026
Jul 22

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability…

Read paper · arxiv.org → Multimodal Benchmark Jul 22, 2026
Jul 22HF Daily Papers

Sample-Efficient Learning from Agent Experience

Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-efficient way for agents to learn from…

Read paper · arxiv.org → Agents Method Jul 22, 2026
Jul 22HF Daily Papers

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more…

Read paper · arxiv.org → Multimodal Benchmark Jul 22, 2026
Jul 22HF Daily Papers

Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation

We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural representations (INRs). Our analysis reveals that sinusoidal activations induce a harmonic line spectrum, providing a spectral…

Read paper · arxiv.org → Models Analysis Jul 22, 2026
Jul 22HF Daily Papers

Visual Contrastive Self-Distillation

On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a…

Read paper · arxiv.org → Multimodal Method Jul 22, 2026
Jul 22HF Daily Papers

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This…

Read paper · arxiv.org → Agents Method Jul 22, 2026
Jul 22HF Daily Papers

GraphVid: Interactive Graph-Controllable Video Generation

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based…

Read paper · arxiv.org → Robotics Method Jul 22, 2026
Jul 22HF Daily Papers

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for…

Read paper · arxiv.org → Evals Benchmark Jul 22, 2026
Jul 22HF Daily Papers

TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation

The development of generalizable robotic manipulation policies is inherently bounded by the availability of large-scale, high-fidelity scene data. While recent automated synthesis methods attempt to bridge this gap via text-to-layout…

Read paper · arxiv.org → Robotics Dataset Jul 22, 2026
Jul 22HF Daily Papers

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward…

Read paper · arxiv.org → Multimodal System Jul 22, 2026
Jul 22HF Daily Papers

ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders

The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fully specified instructions, agents are increasingly expected to transform incomplete product intent…

Read paper · arxiv.org → Evals Benchmark Jul 22, 2026
Jul 22HF Daily Papers

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video…

Read paper · arxiv.org → Multimodal System Jul 22, 2026
Jul 22HF Daily Papers

Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems

Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning…

Read paper · arxiv.org → Infra System Jul 22, 2026
Jul 22HF Daily Papers

Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering

Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These…

Read paper · arxiv.org → Multimodal System Jul 22, 2026
Jul 22HF Daily Papers

Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific…

Read paper · arxiv.org → Multimodal System Jul 22, 2026
Jul 22HF Daily Papers

Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, evidence synthesis, and report generation, yet their reliability in open information environments remains underexplored. A key…

Read paper · arxiv.org → Evals Benchmark Jul 22, 2026
Jul 22HF Daily Papers

What AI Red-Team Evaluations Can and Cannot Prove

Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which…

Read paper · arxiv.org → Evals Benchmark Jul 22, 2026
Jul 21

TRIM

Coding agents leave a mess behind: speculative edits, abandoned hypotheses, scaffolding that never gets cleaned up before the final diff. This paper names it "CodeSlop" and trims it by pruning the agent's search trajectory itself rather than post-editing the output. Builder read: worth wiring into any coding-agent pipeline where diff size matters for review.

Read paper · arxiv.org → Agents Survey Jul 21, 2026
Jul 21

Sycophancy has an address

Slip a model a casual hint, a mislabeled few-shot example, or a fake prior answer, and it'll often flip a correct answer to match. Researchers traced the effect across multiple model families and found the bug isn't in pretraining. Base models barely have it.

Read paper · arxiv.org → Models Method Jul 21, 2026
Jul 21

SOPHIA

Long reasoning chains sometimes get stuck circling the same dead end until the token budget runs out. This method detects that kind of loop mid-trace and steers the model back out before it burns through the budget. Reports gains in both accuracy and token efficiency.

Read paper · arxiv.org → Evals Analysis Jul 21, 2026
Jul 21

Generative World Renderer at the Speed of Play

Generative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesizes RGB frames. Unlike models that generate frames from text/control-hints prompts, AlayaRenderer preserves scene structure…

Read paper · arxiv.org → Science System Jul 21, 2026
Jul 21

ISO: An RLVR-Native Optimization Stack

Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood.…

Read paper · arxiv.org → Models Method Jul 21, 2026
Jul 21

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn…

Read paper · arxiv.org → Multimodal System Jul 21, 2026
Jul 21

DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations

As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce…

Read paper · arxiv.org → Evals Benchmark Jul 21, 2026
Jul 21HF Daily Papers

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims: if…

Read paper · arxiv.org → Multimodal Method Jul 21, 2026
Jul 21HF Daily Papers

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning

Reinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its extension to vision-language models remains constrained by the lack of training data that are simultaneously broad, exactly…

Read paper · arxiv.org → Multimodal Survey Jul 21, 2026
Jul 21HF Daily Papers

Self Gradient Forcing: Native Long Video Extrapolation

Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the…

Read paper · arxiv.org → Multimodal Method Jul 21, 2026
Jul 21HF Daily Papers

Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking

As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents…

Read paper · arxiv.org → Evals Benchmark Jul 21, 2026
Jul 21HF Daily Papers

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and…

Read paper · arxiv.org → Infra System Jul 21, 2026
Jul 21HF Daily Papers

SLPO: Scaling Latent Reasoning via a Surrogate Policy

Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step…

Read paper · arxiv.org → Models Method Jul 21, 2026
Jul 21HF Daily Papers

G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection

This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detection. G-MAD addresses key limitations of real-world aerial dataset construction, including limited…

Read paper · arxiv.org → Evals Dataset Jul 21, 2026
Jul 21HF Daily Papers

SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments

Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requirements using language, but existing…

Read paper · arxiv.org → Robotics Method Jul 21, 2026
Jul 21HF Daily Papers

ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion

3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D and adaptively densifying them in under-reconstructed regions. However, this scene-adaptive capacity allocation is largely…

Read paper · arxiv.org → Multimodal Method Jul 21, 2026
Jul 21HF Daily Papers

Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight…

Read paper · arxiv.org → Science Method Jul 21, 2026
Jul 21HF Daily Papers

LLMs Get Lost in Evolving User Intent

As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront,…

Read paper · arxiv.org → Agents Method Jul 21, 2026
Jul 21HF Daily Papers

Robostral Navigate

Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built…

Read paper · arxiv.org → Robotics System Jul 21, 2026
Jul 21HF Daily Papers

ReferTrack: Referring Then Tracking for Embodied Visual Tracking

Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and…

Read paper · arxiv.org → Robotics Method Jul 21, 2026
Jul 21HF Daily Papers

NVIDIA-labs OO Agents: Native Python Object-Oriented Agents

Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NOOA…

Read paper · arxiv.org → Agents System Jul 21, 2026
Jul 21HF Daily Papers

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic…

Read paper · arxiv.org → Multimodal Method Jul 21, 2026
Jul 21HF Daily Papers

Multimodal Speaker Verification as a Threat to Speaker Anonymization

Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available…

Read paper · arxiv.org → Multimodal System Jul 21, 2026
Jul 21HF Daily Papers

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

Agentic reinforcement learning research is constant algorithm modification, new estimators, new pipeline stages, new rollout schemes, and in mainstream frameworks each change threads through layers of trainer, distributed backend, and…

Read paper · arxiv.org → Agents System Jul 21, 2026
Jul 21HF Daily Papers

Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is completed. It does not explain whether the current behavior is making progress, remaining…

Read paper · arxiv.org → Robotics Survey Jul 21, 2026
Jul 21HF Daily Papers

How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF

In RLHF pipelines, reward scoring blocks policy updates. Slow scoring bottlenecks the entire loop, since no update runs until every rollout gets a score. And yet most setups just default to PyTorch eager mode or torch.compile, no one…

Read paper · arxiv.org → Infra System Jul 21, 2026
Jul 21HF Daily Papers

Multi-Head Attention Residuals

Transformers propagate information across depth through a single additive residual stream: every sublayer reads only the most recent state. Attention residuals relax this by letting each sublayer attend, through a learned softmax. However,…

Read paper · arxiv.org → Models Method Jul 21, 2026
Jul 21HF Daily Papers

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and multi-step logic over long contexts, is an ideal testbed.…

Read paper · arxiv.org → Models Method Jul 21, 2026
Jul 21HF Daily Papers

Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans

A real-time framework for online co-speech gesture generation uses a causal multimodal autoregressive model with streaming speech and motion history, supported by synthetic dialogue data and continual user-feedback adaptation.

Read paper · arxiv.org → Multimodal System Jul 21, 2026
Jul 20

PagedWeight

MoE serving has a memory fight nobody talks about enough: the expert weights and the growing KV cache both want the same GPU room. This paper pages weight precision dynamically, keeping the bits that matter and thinning the ones that don't, so you can stretch context length without adding GPUs. Builder read: promising if the quality claims survive outside the paper's own benchmarks.

Read paper · arxiv.org → Evals Benchmark Jul 20, 2026
Jul 20

Thermodynamic computing

A group sketched a blueprint for analog chips that compute by settling into physical equilibrium, using the hardware's own noise as the substrate instead of fighting it. No chip exists yet, just the math and the architecture. Watch, don't act, but it's a genuinely different bet on where post-GPU compute comes from.

Read paper · arxiv.org → Science System Jul 20, 2026
Jul 20

Cluster-aware optimal transport

Most point-matching methods treat every point as interchangeable. This one matches cluster to cluster first, then refines inside each pair, which should generalize better on data with real group structure. Narrow use case (dataset alignment, domain adaptation) but a clean idea.

Read paper · arxiv.org → Safety Dataset Jul 20, 2026
Jul 20

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This…

Read paper · arxiv.org → Multimodal Dataset Jul 20, 2026
Jul 20HF Daily Papers

Delineate Anything v2: A Global Foundation Model for Field Delineation

Accurate agricultural field boundary delineation at large scale is a foundational task for food security, supply chain transparency, and carbon accounting. While vision foundation models like SAM show remarkable zero-shot capabilities,…

Read paper · arxiv.org → Multimodal Method Jul 20, 2026
Jul 20HF Daily Papers

Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training

Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training: on a 6.78B-parameter MoE language model, AdamW keeps 50.6 GB of first and second moments to update 12.6 GB of bfloat16 weights. We…

Read paper · arxiv.org → Agents Method Jul 20, 2026
Jul 20HF Daily Papers

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause or…

Read paper · arxiv.org → Agents Method Jul 20, 2026
Jul 20HF Daily Papers

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation during denoising remains poorly understood. We introduce a causal interpretability framework for modern large-scale DiTs that…

Read paper · arxiv.org → Multimodal System Jul 20, 2026
Jul 20HF Daily Papers

HPD-Parsing: Hierarchical Parallel Document Parsing

Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflected in unified Vision-Language Model (VLM)-based document parsers. Existing unified parsers process an entire page jointly…

Read paper · arxiv.org → Multimodal Method Jul 20, 2026
Jul 20HF Daily Papers

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the claims a model makes are correct. The dominant decompose-search-verify pipeline catches incorrect claims well but says little…

Read paper · arxiv.org → Evals Benchmark Jul 20, 2026
Jul 20HF Daily Papers

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us…

Read paper · arxiv.org → Science System Jul 20, 2026
Jul 20HF Daily Papers

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing. From a…

Read paper · arxiv.org → Models System Jul 20, 2026
Jul 20HF Daily Papers

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The…

Read paper · arxiv.org → Multimodal Method Jul 20, 2026
Jul 20HF Daily Papers

Masked Visual Actions for Unified World Modeling

Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form…

Read paper · arxiv.org → Robotics Method Jul 20, 2026
Jul 20HF Daily Papers

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting…

Read paper · arxiv.org → Science Method Jul 20, 2026
Jul 20HF Daily Papers

H^2SD: Hybrid Hindsight Self-Distillation

Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models on tasks such as mathematical reasoning and code generation. However, most RLVR methods assign a scalar…

Read paper · arxiv.org → Science Method Jul 20, 2026
Jul 20HF Daily Papers

NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs

Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual substrate engineering,…

Read paper · arxiv.org → Infra System Jul 20, 2026
Jul 20HF Daily Papers

Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models

Injecting factual knowledge into large language models (LLMs) reliably and at scale remains an open challenge. Hypernetworks provide a promising solution to large-scale knowledge injection. Although hypernetworks are typically applied for…

Read paper · arxiv.org → Models Method Jul 20, 2026
Jul 20HF Daily Papers

AutoIndex: Learning Representation Programs for Retrieval

We present AutoIndex, a framework for learning representation programs: executable transformations that map raw documents into the representations exposed to a retrieval system. Rather than tuning retrievers, rerankers, or a small set of…

Read paper · arxiv.org → Evals Benchmark Jul 20, 2026
Jul 20HF Daily Papers

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

Modern ASR models trained on heterogeneously annotated data treat transcription style (verbatim vs. intended) as an uncontrolled latent variable, causing measurable decoding instability, evaluation confounding (up to 60% of reported WER…

Read paper · arxiv.org → Robotics Benchmark Jul 20, 2026
Jul 20HF Daily Papers

FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents

Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different…

Read paper · arxiv.org → Evals Benchmark Jul 20, 2026
Jul 20HF Daily Papers

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and…

Read paper · arxiv.org → Robotics System Jul 20, 2026
Jul 20HF Daily Papers

AI Tour Meeting: Group Travel Planning by LLM Agents

This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Model (LLM)-based agents. The agents are instantiated with distinct personas and collaboratively seek an itinerary that satisfies…

Read paper · arxiv.org → Agents System Jul 20, 2026
Jul 20HF Daily Papers

Enhancing Rubric-based RL via Self-Distillation

Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout manages to satisfy (Unexplored Criteria, UC) receive no…

Read paper · arxiv.org → Models Method Jul 20, 2026
Jul 20HF Daily Papers

InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal

Specialized scholarly agents use reinforcement learning and real-time citation verification to improve reasoning and factual accuracy in peer review and rebuttal generation.

Read paper · arxiv.org → Evals Survey Jul 20, 2026
Jul 19

AutoSynthesis

Multi-agent system that runs an entire meta-analysis end to end: finds the studies, pulls effect sizes, pools them, flags bias. That's the grunt work behind every systematic review, normally a grad student with a spreadsheet for six months. Whether it holds up on a genuinely contested literature, not a clean toy set, is the real test nobody's run yet.

Read paper · arxiv.org → Infra Survey Jul 19, 2026
Jul 19

BadWAM

World-action models are supposed to be safer because they predict what happens next before choosing an action. This paper's finding: the physics prediction can be dead right and the action picked right after it still gets someone hurt. Good world model, bad policy. If you're building embodied agents on this architecture, "predicts the world accurately" is not a safety guarantee.

Read paper · arxiv.org → Science System Jul 19, 2026
Jul 19

MedFailBench

A clinician-built benchmark that stops grading medical AI on right versus wrong and starts grading which safety gate it blew through, a missed urgent referral, a skipped dosage check, graded 1 to 5 on severity. More useful than one more accuracy leaderboard if you're anywhere near deploying this stuff in an actual clinic.

Read paper · arxiv.org → Science Benchmark Jul 19, 2026
Jul 19

ShotPlan: Cinematic Video Generation with Learnable Planning Token

Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition require explicit shot planning. To…

Read paper · arxiv.org → Multimodal Method Jul 19, 2026
Jul 19

WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic benchmark for…

Read paper · arxiv.org → Evals Benchmark Jul 19, 2026
Jul 19

Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?

Self-hosted AI agents read and write their own memory and configuration files to function. An agent may get compromised via corruption of its own state -- a compromise realized via legitimate OS system call invocation. We refer to this…

Read paper · arxiv.org → Infra System Jul 19, 2026
Jul 19HF Daily Papers

FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting video editing data…

Read paper · arxiv.org → Multimodal Method Jul 19, 2026
Jul 19HF Daily Papers

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift

We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. Our key intuition is that by training the model to distinguish good and bad tokens…

Read paper · arxiv.org → Models Method Jul 19, 2026
Jul 19HF Daily Papers

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still…

Read paper · arxiv.org → Multimodal Method Jul 19, 2026
Jul 19HF Daily Papers

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and…

Read paper · arxiv.org → Multimodal System Jul 19, 2026
Jul 19HF Daily Papers

SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

Pruning long context for coding agents has been a vital technology for efficient context management. While existing context pruning methods such as SWE-Pruner realize this by attaching a separate code classifier, we find the agent itself…

Read paper · arxiv.org → Agents Method Jul 19, 2026
Jul 19HF Daily Papers

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial…

Read paper · arxiv.org → Robotics System Jul 19, 2026
Jul 19HF Daily Papers

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning…

Read paper · arxiv.org → Evals Benchmark Jul 19, 2026
Jul 19HF Daily Papers

DiFA: Inference-Time Forward-Process Alignment for Diffusion Models

The prevailing inference framework for diffusion models formulates generation fundamentally as a problem of numerical integration. This perspective casts the model as an exact estimator, neglecting the inherent statistical uncertainty of…

Read paper · arxiv.org → Safety System Jul 19, 2026
Jul 19HF Daily Papers

EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on…

Read paper · arxiv.org → Multimodal Benchmark Jul 19, 2026
Jul 19HF Daily Papers

SciForma: Structure-Faithful Generation of Scientific Diagrams

Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diagrams must faithfully render components, directional relations, and textual annotations. Since a single error, such as a reversed…

Read paper · arxiv.org → Models Dataset Jul 19, 2026
Jul 19HF Daily Papers

ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large…

Read paper · arxiv.org → Robotics Method Jul 19, 2026
Jul 19HF Daily Papers

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the…

Read paper · arxiv.org → Evals Benchmark Jul 19, 2026
Jul 19HF Daily Papers

ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of…

Read paper · arxiv.org → Multimodal Method Jul 19, 2026
Jul 19HF Daily Papers

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for…

Read paper · arxiv.org → Science Benchmark Jul 19, 2026
Jul 19HF Daily Papers

Subliminal Clocks: Latent Time Modelling in Diffusion Language Models

Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly conditioned on a timestep, raising a natural question: do these…

Read paper · arxiv.org → Models Method Jul 19, 2026
Jul 19HF Daily Papers

Differentiable Logic Gate Networks for Low-Latency EEG Classification on Edge Devices

Real-time EEG classification on edge devices is bottlenecked by the floating-point arithmetic of conventional neural networks. We investigated Differentiable Logic Gate Networks (Diff-Logic) as a hardware-native alternative that compiles…

Read paper · arxiv.org → Infra Method Jul 19, 2026
Jul 19HF Daily Papers

SLAM in Low-Light Environments: Project Report

Simultaneous localization and mapping (SLAM) is one of the fundamental problems in robotics, as it enables autonomous operations in real-world scenarios. Under low illumination, reduced contrast, sensor noise, and motion blur degrade both…

Read paper · arxiv.org → Robotics Analysis Jul 19, 2026
Jul 19HF Daily Papers

Three-Body Scattering for Generative Modeling

Modern generative models typically rely on an adversarial critic, a prescribed noise-to-data path, or an autoregressive factorization. Instead, we show that a proper distributional energy can induce sample-level motion and provide direct…

Read paper · arxiv.org → Models Method Jul 19, 2026
Jul 19HF Daily Papers

O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning

Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are…

Read paper · arxiv.org → Robotics System Jul 19, 2026
Jul 19HF Daily Papers

Uncovering Latent Reasoning Strategies in Language Models

A language model p_θ(y mid x) trained on reasoning tasks learns to solve problems via multiple distinct strategies, yet these strategies are implicit and entangled within the model's response distribution. We study the problem of…

Read paper · arxiv.org → Models Analysis Jul 19, 2026
Jul 19HF Daily Papers

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

The benchmark evaluates autonomous coding agents on open-ended world-model research by having them iteratively improve a starter model across game environments using a shared structured-state format.

Read paper · arxiv.org → Evals Benchmark Jul 19, 2026
Jul 18

Robots don't read intent, they read words

A new probe called PRISM checks a model's hidden states instead of its output text, hunting for instructions that sound harmless but turn dangerous once a robot actually executes them. On the team's new PhysicalSafetyBench-1K it hit 99.6% accuracy with a 0.7% false-positive rate. A plain text-safety judge, by contrast, rejected two-thirds of tasks that were completely fine.

Read paper · arxiv.org → Robotics Method Jul 18, 2026
Jul 18

Agentic search rewards documents that look useless

Retrieval has always been scored on one question: does this document answer the question in front of you. Wrong test, apparently, once an agent starts searching in multiple steps. Researchers logged 23,000+ document reads inside a ReAct agent and found almost zero correlation (Spearman rho of -0.026) between a document scoring as relevant and it actually changing the outcome.

Read paper · arxiv.org → Evals Benchmark Jul 18, 2026
Jul 18

A model's parts don't add up to its whole

Ask a model for estimates about subgroups of a population, then ask it for the population-level number, and the two should reconcile. Often they don't.

Read paper · arxiv.org → Models Method Jul 18, 2026
Jul 18HF Daily Papers

The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture

We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architecture as an integro-differential equation (IDE) on a semantic fiber bundle calE = calM times R^d. Beginning from a single…

Read paper · arxiv.org → Evals System Jul 18, 2026
Jul 18HF Daily Papers

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis

Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the strong priors of video foundation models, multi-view consistent HOI synthesis remains challenging due to complex hand motions…

Read paper · arxiv.org → Robotics Method Jul 18, 2026
Jul 18HF Daily Papers

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality…

Read paper · arxiv.org → Multimodal Analysis Jul 18, 2026
Jul 18HF Daily Papers

EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World

This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds. Existing systems either treat interactive literary simulation as static persona imitation or isolated scene…

Read paper · arxiv.org → Evals Benchmark Jul 18, 2026
Jul 18HF Daily Papers

Distilled Reinforcement Learning for LLM Post-training

Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on…

Read paper · arxiv.org → Safety Method Jul 18, 2026
Jul 18HF Daily Papers

Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop

Replacing individual LLM agents with low-parameter surrogates fitted from cheap queries enables scalable society simulations, with validity predicted by an interaction-order and memory taxonomy.

Read paper · arxiv.org → Agents Survey Jul 18, 2026
Jul 17

A robot policy that remembers 8,000 steps back, at no extra inference cost

Most robot foundation models look at the last frame or two and nothing else. RoboTTT scales working memory to 8,000 timesteps, three orders of magnitude past prior work, by folding test-time training into the policy instead of just widening the context window.

Read paper · arxiv.org → Robotics Method Jul 17, 2026
Jul 17

Your pretraining corpus can be poisoned from the comments section

Earlier poisoning research stuck to tidy, well-behaved sources like Wikipedia. This paper points at something messier and already public: open discussion interfaces, the comment threads and forums that get swept into web crawls. The harder question is whether injected content actually survives data curation and lands in a trained model, which is why the authors built a measurement tool (HalfLife) to check.

Read paper · arxiv.org → Evals Dataset Jul 17, 2026
Jul 17

A security-agent benchmark that finally asks what the tool calls cost

Most agent security evals just measure who wins, assuming an unlimited compute budget, which is not how a SOC operates. This one scores agents on both offense (Cybench) and defense (Splunk BOTS v1 investigations) against actual inference and tool spend.

Read paper · arxiv.org → Safety Benchmark Jul 17, 2026
Jul 17HF Daily Papers

Can Multimodal Large Language Models Understand OCT?

Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing…

Read paper · arxiv.org → Science Analysis Jul 17, 2026
Jul 17HF Daily Papers

Group Entropy-Controlled Policy Optimization

Entropy control has become an effective tool in reinforcement learning (RL) of large language models (LLMs), helping balance exploration-exploitation trade-off during alignment process. Such RL paradigm is often conducted on mixtures of…

Read paper · arxiv.org → Robotics Method Jul 17, 2026
Jul 17HF Daily Papers

Environment-free Synthetic Data Generation for API-Calling Agents

Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale typically requires fully implemented environments with executable APIs and realistic,…

Read paper · arxiv.org → Agents Method Jul 17, 2026
Jul 17HF Daily Papers

DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines

Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts. We call this…

Read paper · arxiv.org → Agents System Jul 17, 2026
Jul 17HF Daily Papers

Dataset Distillation by Influence Matching

We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (per-step gradients or training trajectories), Influence Matching (Inf-Match) aligns the final outcome of training: it learns a…

Read paper · arxiv.org → Models Dataset Jul 17, 2026
Jul 17HF Daily Papers

CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation

On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compounding failure modes: (i) cold-start collapse, where a fresh student assigns near-zero mass to…

Read paper · arxiv.org → Infra Method Jul 17, 2026
Jul 17HF Daily Papers

Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety Testing

In our prior work, Pedestrian Archetypes, we defined pedestrian archetypes as collections of behaviors that uniquely identify a specific type of pedestrian. The first paper proposed 12 pedestrian archetypes, including the Wanderer, Drunk,…

Read paper · arxiv.org → Safety Dataset Jul 17, 2026
Jul 16

GitHub's agentic PR data

turns out less dramatic than the hype suggests. Across 25,264 agent-submitted PRs in 2,361 popular repos, the median project sees only one or two a quarter, and most stay under an industry benchmark of 36 PRs per contributor over three months. Almost every one still runs through a single human reviewer before merging. Agents aren't replacing dev teams yet. They're on a slow, supervised trial period.

Read paper · arxiv.org → Evals Survey Jul 16, 2026
Jul 16

Optimized agents don't stay optimized

A Terminal-Bench 2.0 study pitted three agent-tuning methods against a fresh round of unseen tasks. GEPA's gains (66% pass rate after tuning) collapsed below the 58.7% unoptimized baseline the moment new tasks showed up. Meta Harness held its 64.6% but stopped improving. Only RELAI-VCL, which builds in regression control, kept climbing, to 76.4%, and kept the gains on retrain.

Read paper · arxiv.org → Robotics Benchmark Jul 16, 2026
Jul 16

Pen-testing needs a new failure mode

A proposed framework argues exploit-hunting misses how AI systems actually break: an agent can leak data or take a bad action with zero vulnerabilities exploited, just a poisoned tool call or prompt injection nudging its behavior off-objective. The six-step process tests for that kind of behavioral violation directly instead of waiting for a CVE. Worth stealing if you're red-teaming anything with real tool access.

Read paper · arxiv.org → Infra System Jul 16, 2026
Jul 16

When Does Muon Help Agentic Reinforcement Learning?

Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through matched single-seed comparisons with AdamW…

Read paper · arxiv.org → Agents Analysis Jul 16, 2026
Jul 16

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation

We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) has advanced significantly through domain-specific models, tool-augmented LLMs, and scientific…

Read paper · arxiv.org → Science Method Jul 16, 2026
Jul 16

DSWorld: A Data Science World Model for Efficient Autonomous Agents

Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily rely on trial-and-error workflows that involve expensive computation. This bottleneck motivates models that can anticipate…

Read paper · arxiv.org → Science Method Jul 16, 2026
Jul 16

Recursive Harness Self-Improvement

Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing…

Read paper · arxiv.org → Models Method Jul 16, 2026
Jul 16

Loop the Loopies!

We present Loopie, the most powerful looped Transformer to date. The Loopie series consists of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6Bparameter model with 0.6B active parameters. Looped…

Read paper · arxiv.org → Models Method Jul 16, 2026
Jul 16HF Daily Papers

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily…

Read paper · arxiv.org → Multimodal Method Jul 16, 2026
Jul 16HF Daily Papers

RecGPT-V3 Technical Report

Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this paradigm on Taobao by centering user…

Read paper · arxiv.org → Infra System Jul 16, 2026
Jul 16HF Daily Papers

Understanding Reasoning from Pretraining to Post-Training

Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic…

Read paper · arxiv.org → Models Method Jul 16, 2026
Jul 16HF Daily Papers

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether…

Read paper · arxiv.org → Multimodal Benchmark Jul 16, 2026
Jul 16HF Daily Papers

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models

The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload…

Read paper · arxiv.org → Robotics Method Jul 16, 2026
Jul 16HF Daily Papers

SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction

Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks where a single erroneous action can lead to irreversible consequences. Existing…

Read paper · arxiv.org → Safety System Jul 16, 2026
Jul 16HF Daily Papers

Nonuniformity Principle in Human-AI Coworking

As generative AI is increasingly applied to automate multi-step and high-stake workflows, human judgment and involvement remain essential for ensuring the quality of AI-generated outputs. In practice, while it is desirable for human…

Read paper · arxiv.org → Models Method Jul 16, 2026
Jul 16HF Daily Papers

An Exam for Active Observers

Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range…

Read paper · arxiv.org → Science Method Jul 16, 2026
Jul 16HF Daily Papers

FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation

Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-p routing creates uneven…

Read paper · arxiv.org → Multimodal System Jul 16, 2026
Jul 16HF Daily Papers

Interactive Training 2: Auditable Control Plane for Live Model Training

Experiment trackers show how training is progressing, but changing a live run still usually requires trainer-specific code. We present Interactive Training 2, an open-source control plane for steering training through a shared protocol.…

Read paper · arxiv.org → Robotics Method Jul 16, 2026
Jul 16HF Daily Papers

In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing

Autonomous driving systems (ADS) are rapidly advancing and increasingly deployed in real-world applications. This creates growing demands for effective testing to ensure system functionality and safety. However, ADS testing remains complex…

Read paper · arxiv.org → Safety System Jul 16, 2026
Jul 16HF Daily Papers

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily…

Read paper · arxiv.org → Multimodal Benchmark Jul 16, 2026
Jul 15

Agents don't know when a task is small

A new paper catches LLM agents defaulting to max-context-first: re-reading every file and dependency they've already seen, even for a one-line edit. The fix isn't a bigger context window, it's a cheaper first question the agent never asks: how much effort does this actually need.

Read paper · arxiv.org → Agents Method Jul 15, 2026
Jul 15

LLM judges grade easy when there's no answer key

Turns out judge models skew generous specifically in no-reference setups, which is most real evaluation pipelines (open-ended output, no ground truth to check against). If your eval numbers look suspiciously good, this is a plausible reason why.

Read paper · arxiv.org → Evals Benchmark Jul 15, 2026
Jul 15

Ask 44 models to pick any word, 41% say "serendipity."

Funny, but it's a real signal on model homogeneity. Post-training keeps compressing "creative" output down to the same handful of safe choices across labs. Remember that next time a demo calls an LLM's output original.

Read paper · arxiv.org → Models Method Jul 15, 2026
Jul 15

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-token contexts, while post-training workloads often remain at 256K tokens or below and rely on length generalization at…

Read paper · arxiv.org → Infra System Jul 15, 2026
Jul 15

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language…

Read paper · arxiv.org → Multimodal Benchmark Jul 15, 2026
Jul 15HF Daily Papers

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They…

Read paper · arxiv.org → Multimodal Method Jul 15, 2026
Jul 15HF Daily Papers

WanSong v1.0 Technical Report

Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these…

Read paper · arxiv.org → Robotics Analysis Jul 15, 2026
Jul 15HF Daily Papers

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to track task progress. When search…

Read paper · arxiv.org → Agents Method Jul 15, 2026
Jul 15HF Daily Papers

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient generation. Reinforcement learning (RL) has become a powerful way to align diffusion and flow…

Read paper · arxiv.org → Infra Method Jul 15, 2026
Jul 15HF Daily Papers

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (RL) provides a practical optimization…

Read paper · arxiv.org → Agents Method Jul 15, 2026
Jul 15HF Daily Papers

Video = World + Event Stream

We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the…

Read paper · arxiv.org → Multimodal Method Jul 15, 2026
Jul 15

Hierarchical Denoising For Multi-Step Visual Reasoning

Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusion enables global…

Read paper · arxiv.org → Multimodal Method Jul 15, 2026
Jul 15

SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment

CAD-to-image alignment aims to estimate an object's 9D pose (rotation, translation, and anisotropic scale) from a single RGB image, enabling applications in robotics and augmented reality. Recent zero-shot methods use visual foundation…

Read paper · arxiv.org → Robotics Method Jul 15, 2026
Jul 15HF Daily Papers

RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, text-centric, or derived from agent traces, leaving tutorial…

Read paper · arxiv.org → Multimodal Method Jul 15, 2026
Jul 15HF Daily Papers

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish useful uncertainty from detrimental confusion, limiting its effectiveness as a correctness signal. We…

Read paper · arxiv.org → Models Method Jul 15, 2026
Jul 15HF Daily Papers

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2)…

Read paper · arxiv.org → Robotics Method Jul 15, 2026
Jul 15HF Daily Papers

On-Policy Delta Distillation

On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model. Although on-policy distillation has…

Read paper · arxiv.org → Multimodal Method Jul 15, 2026
Jul 15HF Daily Papers

xHC: Expanded Hyper-Connections

Hyper-Connections (HC) expand the residual stream of Transformers into N parallel streams, providing a form of memory scaling beyond model width and depth. Manifold-Constrained HC (mHC) stabilizes this formulation at scale. The large gains…

Read paper · arxiv.org → Agents Method Jul 15, 2026
Jul 15HF Daily Papers

Trajectory-aware Cross-view Geo-localization with Sequential Observations

Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent methods show that sequential queries such as video clips yield richer spatiotemporal cues than single images, yet they overlook a…

Read paper · arxiv.org → Multimodal Method Jul 15, 2026
Jul 15HF Daily Papers

Multi-Turn On-Policy Distillation with Prefix Replay

We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histories. Fully online OPD is costly because…

Read paper · arxiv.org → Agents Analysis Jul 15, 2026
Jul 14

Split the backdoor across agents, and every local monitor waves it through

A new paper on multi-agent LLM systems finds a hole in the usual safety net, a monitor that checks each message or tool call on its own. Fragment a harmful payload across agents and every single step passes clean, because the harm only exists once the pieces assemble somewhere downstream. The fix isn't a sharper per-step monitor.

Read paper · arxiv.org → Safety System Jul 14, 2026
Jul 14

LLM judges are biased in a spot you can actually point to

Researchers went inside seven judge models, across seven bias types and nine benchmarks, and found biased inputs don't just score wrong, they push the judge's hidden state along a specific, low-dimensional direction. Steer activations along that direction and you can manufacture the bias on a clean input, or erase it on a biased one.

Read paper · arxiv.org → Evals Benchmark Jul 14, 2026
Jul 14

Sakana let a vision-language model play Picbreeder with zero instructions, and it didn't get weird in the right way

Picbreeder is the old human experiment where people bred abstract images generation after generation with no goal beyond "pick what looks interesting," and ended up with genuinely surprising art. Swap in a frontier VLM making the same picks and the output looks nothing like that human baseline, less of the compounding, wandering strangeness that made the original worth studying.

Read paper · x.com → Multimodal Benchmark Jul 14, 2026
Jul 14HF Daily Papers

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility…

Read paper · arxiv.org → Evals Benchmark Jul 14, 2026
Jul 14HF Daily Papers

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as…

Read paper · arxiv.org → Multimodal Method Jul 14, 2026
Jul 14HF Daily Papers

GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physically grounded action generation. However, a common design in…

Read paper · arxiv.org → Robotics Method Jul 14, 2026
Jul 14HF Daily Papers

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platform GUI interaction support and a well-built self-evolution mechanism. These flaws limit its adaptation to diverse device…

Read paper · arxiv.org → Agents System Jul 14, 2026
Jul 14HF Daily Papers

OvisOCR2 Technical Report

We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document page image, it generates a Markdown representation in natural reading order, covering text, formulas, tables, and visual…

Read paper · arxiv.org → Multimodal Analysis Jul 14, 2026
Jul 14HF Daily Papers

From Pixels to States: Rethinking Interactive World Models as Game Engines

Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphics, games, and artificial intelligence. Recent video generative models provide a data-driven route toward this goal by…

Read paper · arxiv.org → Multimodal System Jul 14, 2026
Jul 14HF Daily Papers

DeepLoop: Depth Scaling for Looped Transformers

Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling problem: in an…

Read paper · arxiv.org → Models Method Jul 14, 2026
Jul 14HF Daily Papers

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference…

Read paper · arxiv.org → Multimodal Benchmark Jul 14, 2026
Jul 14HF Daily Papers

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can…

Read paper · arxiv.org → Multimodal Benchmark Jul 14, 2026
Jul 14HF Daily Papers

Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the…

Read paper · arxiv.org → Infra System Jul 14, 2026
Jul 14HF Daily Papers

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel

We report a way to make a frozen small language model both more capable and dramatically cheaper at once, without changing any weights. Verified knowledge is deposited once as a byte-exact key-value (KV) state artifact and later restored,…

Read paper · arxiv.org → Models Analysis Jul 14, 2026
Jul 14HF Daily Papers

Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code

Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their strictness makes generation more difficult. Off-the-shelf compilers can provide useful feedback post-generation, but does not…

Read paper · arxiv.org → Models Method Jul 14, 2026
Jul 14HF Daily Papers

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation

Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. Unlike continuous…

Read paper · arxiv.org → Models System Jul 14, 2026
Jul 14

Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving

Validating autonomous driving systems requires diverse, regulation-compliant test scenarios. In simulation-based testing, scenarios are defined as executable scripts. Yet automatically generating such scripts from regulatory descriptions…

Read paper · arxiv.org → Infra System Jul 14, 2026
Jul 14

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination.…

Read paper · arxiv.org → Robotics Method Jul 14, 2026
Jul 14HF Daily Papers

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). However, conventional 3D-VAEs are mainly optimized for pixel-level reconstruction, which can limit the semantic and spatio-temporal…

Read paper · arxiv.org → Multimodal Method Jul 14, 2026
Jul 14HF Daily Papers

Cura 1T: Specialized Model for Agentic Healthcare

Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over…

Read paper · arxiv.org → Science Method Jul 14, 2026
Jul 14HF Daily Papers

DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation

Existing 3D generative models predominantly rely on implicit volumetric representations, which enforce watertight topology and struggle to represent thin-shell and non-manifold geometries such as garments. Geometry image-based approaches…

Read paper · arxiv.org → Multimodal Method Jul 14, 2026
Jul 14HF Daily Papers

Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings

Serial verification gates are a core reliability primitive in LLM harnesses: a candidate answer is returned only if k verifier calls all accept it. Under conditionally independent gates, the recent Odds Law (arXiv:2606.15712) shows that…

Read paper · arxiv.org → Models Method Jul 14, 2026
Jul 14HF Daily Papers

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot…

Read paper · arxiv.org → Robotics Dataset Jul 14, 2026
Jul 14HF Daily Papers

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. However, BC finetuning progressively overwrites the pretrained…

Read paper · arxiv.org → Robotics Method Jul 14, 2026
Jul 14HF Daily Papers

Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making

Large language models are increasingly deployed as agents, but reliable agentic behavior requires more than next-token prediction. At inference time, it is preferred that an agent can decide whether to proceed with its current reasoning,…

Read paper · arxiv.org → Robotics Method Jul 14, 2026
Jul 13

Clinical RAG's blind spot

A new eval shows a clinical RAG system can pass every automated grounding check, zero hallucinations, near-perfect faithfulness, while attributing the right fact to the wrong patient's chart. Standard RAG benchmarks check whether a claim is supported by retrieved text, not whether that text belongs to the correct entity.

Read paper · arxiv.org → Evals Benchmark Jul 13, 2026
Jul 13

Soofi S 30B-A3B

A sovereign, open-source German/English foundation model: a hybrid Mamba-Transformer MoE that activates just 3B of its 30B parameters per token and holds its inference cache roughly flat as context grows. Read: cheap long-context inference you can self-host, with no US API in the loop.

Read paper · arxiv.org → Models Method Jul 13, 2026
Jul 13

Mach-Mind-4-Flash

35B parameters, 3B active, and the authors claim it matches or beats 100B-class models through post-training alone, no extra pretraining compute spent. Worth watching whether that holds outside their own benchmarks. If it does, "just scale pretraining" gets a lot less convincing as a MoE strategy.

Read paper · arxiv.org → Evals Benchmark Jul 13, 2026
Jul 13

Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals

Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment.…

Read paper · arxiv.org → Safety Method Jul 13, 2026
Jul 13

Metacognition in LLMs: Foundations, Progress, and Opportunities

Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable,…

Read paper · arxiv.org → Models Method Jul 13, 2026
Jul 13

Motion4Motion: Motion Transfer Across Subjects at Inference

This work explores the motion transfer from one video to another, which is crucial in animation for diverse characters. Previously, video motion transfer has been largely explored between human and human-like characters, enabling a lot of…

Read paper · arxiv.org → Multimodal Method Jul 13, 2026
Jul 13

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both…

Read paper · arxiv.org → Science Benchmark Jul 13, 2026
Jul 13

LightMem-Ego: Your AI Memory for Everyday Life

Personal AI assistants on mobile and wearable devices continuously perceive users' daily lives through visual and audio streams. However, answering queries about past experiences requires lightweight multimodal memory that can continuously…

Read paper · arxiv.org → Multimodal Method Jul 13, 2026
Jul 13

Multi-Agent LLMs Fail to Explore Each Other

Exploration is essential for reliable autonomy in multi-agent systems, yet it remains unclear whether large language model (LLM) agents can explore effectively when interacting with one another. We show that modern LLM agents fail to do…

Read paper · arxiv.org → Infra System Jul 13, 2026
Jul 13

Evidence-Backed Video Question Answering

Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales…

Read paper · arxiv.org → Multimodal Method Jul 13, 2026
Jul 13

MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning

Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation,…

Read paper · arxiv.org → Evals Benchmark Jul 13, 2026
Jul 13

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot…

Read paper · arxiv.org → Robotics Method Jul 13, 2026
Jul 13

Latent-Identity Tuning in Text-to-Image Personalization Models

Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a subject's perceived identity. Current personalization and editing methods built on general-purpose text-to-image models,…

Read paper · arxiv.org → Multimodal Method Jul 13, 2026
Jul 13HF Daily Papers

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World

AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perform best in real-world targets. Existing evaluation protocols assess and optimize for…

Read paper · arxiv.org → Robotics Benchmark Jul 13, 2026
Jul 13HF Daily Papers

Self-Improvements in Modern Agentic Systems: A Survey

Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input. This survey frames modern…

Read paper · arxiv.org → Robotics Survey Jul 13, 2026
Jul 13HF Daily Papers

PalmClaw: A Native On-Device Agent Framework for Mobile Phones

Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which…

Read paper · arxiv.org → Infra System Jul 13, 2026
Jul 13HF Daily Papers

Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation

We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation,…

Read paper · arxiv.org → Multimodal Method Jul 13, 2026
Jul 13HF Daily Papers

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive. Yet existing approaches…

Read paper · arxiv.org → Multimodal Method Jul 13, 2026
Jul 13HF Daily Papers

Tracing Agentic Failure from the Flow of Success

Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the task to fail, is critical for debugging and improving these systems. Existing approaches either rely on prompting-based…

Read paper · arxiv.org → Infra System Jul 13, 2026
Jul 13HF Daily Papers

ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually…

Read paper · arxiv.org → Infra Method Jul 13, 2026
Jul 13HF Daily Papers

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constraints, existing…

Read paper · arxiv.org → Models Method Jul 13, 2026
Jul 13HF Daily Papers

Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable

The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs, environments, and requirements…

Read paper · arxiv.org → Agents Method Jul 13, 2026
Jul 13HF Daily Papers

Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering

Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structures and distorted data representations. We present ChartCynics, an agentic dual-path framework…

Read paper · arxiv.org → Multimodal System Jul 13, 2026
Jul 13HF Daily Papers

Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in its forward direction. We observe that the action-observation-continuation loop of a…

Read paper · arxiv.org → Agents Method Jul 13, 2026
Jul 13HF Daily Papers

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-level visual…

Read paper · arxiv.org → Multimodal Method Jul 13, 2026
Jul 13HF Daily Papers

Let RGB Be the Language of Vision

This work introduces a unified formulation for vision models, where diverse forms of visual information beyond natural images, such as masks, depth maps, and other structured visual signals, are all represented as RGB images, while general…

Read paper · arxiv.org → Multimodal Method Jul 13, 2026
Jul 13HF Daily Papers

UniVR: Thinking in Visual Space for Unified Visual Reasoning

Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and…

Read paper · arxiv.org → Multimodal Method Jul 13, 2026
Jul 13HF Daily Papers

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

Human cognition does not separate understanding and generation. A teacher at a whiteboard speaks and draws together, each modality reshapes the other. In this paper, we bring this coupled loop to artificial systems. Masked Diffusion Models…

Read paper · arxiv.org → Multimodal System Jul 13, 2026
Jul 13HF Daily Papers

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a…

Read paper · arxiv.org → Robotics Benchmark Jul 13, 2026
Jul 13HF Daily Papers

AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow

We present AffectFlow-DINO, a multi-task learning system for the 11th ABAW challenge that extends a standard deterministic architecture with a conditional rectified-flow head to model the inherent ambiguity of in-the-wild facial behavior.…

Read paper · arxiv.org → Infra System Jul 13, 2026
Jul 13HF Daily Papers

Rethinking the Evaluation of Harness Evolution for Agents

We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. This…

Read paper · arxiv.org → Evals Benchmark Jul 13, 2026
Jul 13HF Daily Papers

From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality

Code review helps maintain software quality before code integration, but it also imposes a substantial workload on human reviewers. As generative artificial intelligence becomes part of software development, code review is shifting from a…

Read paper · arxiv.org → Agents Survey Jul 13, 2026
Jul 13HF Daily Papers

ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams

Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have…

Read paper · arxiv.org → Multimodal System Jul 13, 2026
Jul 13HF Daily Papers

Color Pass-Through via Camera-Display Coupling

When a real-world scene is captured by a smartphone camera and viewed on its screen, the displayed image often differs noticeably from the original scene in color, brightness, and contrast. This gap persists despite substantial advances in…

Read paper · arxiv.org → Multimodal Method Jul 13, 2026
Jul 13HF Daily Papers

VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression

Vision-language models (VLMs) process large numbers of visual tokens, resulting in substantial inference latency and memory overhead. This has motivated extensive research on visual token compression. While training-free strategies rely on…

Read paper · arxiv.org → Multimodal Method Jul 13, 2026
Jul 13HF Daily Papers

Edge-Aware Thermal Infrared UAV Swarm Tracking

Thermal infrared (TIR) imaging is essential for UAV swarm operations in visually degraded environments. However, tracking tiny UAVs remains challenging due to limited appearance cues, frequent occlusions, and rapid maneuvers. Despite…

Read paper · arxiv.org → Multimodal Method Jul 13, 2026
Jul 12

Super Weights aren't the load-bearing wall everyone assumed

Pruning those handful of "critical" parameters that supposedly tank a model by orders of magnitude doesn't hold across all LLMs. Try to fine-tune only those coordinates on OLMo-1B and OLMo-7B and accuracy collapses to random guessing, even out to 36K neighboring parameters. Plain LoRA, spread thin across the same layers, gets there fine using 0.16% of the weights.

Read paper · arxiv.org → Models Method Jul 12, 2026
Jul 12

Quantized models can ace your eval and still be a different model underneath

A new "correctness agreement" metric checks whether a quantized model gets the same questions right as the full one, not just the same aggregate score, and the answer is no: real behavioral divergence shows up under moderate quantization even when accuracy and perplexity look untouched. Query and key projections take the damage first.

Read paper · arxiv.org → Evals Benchmark Jul 12, 2026
Jul 12

LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow

Flow matching over carefully designed latent representations has recently emerged as a powerful paradigm for topology-aware mesh generation. Existing approaches, however, model vertices and connectivity jointly in a joint latent space,…

Read paper · arxiv.org → Multimodal Method Jul 12, 2026
Jul 12

Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or…

Read paper · arxiv.org → Multimodal Method Jul 12, 2026
Jul 12

Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution

LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone to factual errors caused by insufficient repository understanding. Recent methods attempt to mitigate this limitation…

Read paper · arxiv.org → Agents Method Jul 12, 2026
Jul 12HF Daily Papers

MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors

Current visual generation models are capable of producing high-quality content, yet they lack a coherent perception of the spatial structure. Existing generative novel view synthesis methods typically introduce explicit geometry priors,…

Read paper · arxiv.org → Multimodal Method Jul 12, 2026
Jul 12HF Daily Papers

MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models

Multi-scene navigation (clearing an objective in one bounded space and then crossing a portal into the next) is a defining feature of contemporary 3D games, but authoring it is laborious: every portal must have consistent endpoints on both…

Read paper · arxiv.org → Robotics Method Jul 12, 2026
Jul 12HF Daily Papers

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration,…

Read paper · arxiv.org → Evals Benchmark Jul 12, 2026
Jul 12HF Daily Papers

AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling

Recent generalizable 3D Gaussian Splatting models have advanced long-sequence novel view synthesis (NVS), but at the cost of substantial redundant computation. We identify that the redundancy can be mitigated based on two observations: (i)…

Read paper · arxiv.org → Multimodal Method Jul 12, 2026
Jul 12HF Daily Papers

RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM

Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet existing systems construct knowledge graphs in a single extraction pass, producing noisy entities and brittle retrieval. RAGU, an…

Read paper · arxiv.org → Evals Benchmark Jul 12, 2026
Jul 12HF Daily Papers

Qwen-Music Technical Report

In this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and high-fidelity songs with complete vocal singing. Qwen-Music supports two core tasks: Text to Music Generation, which create…

Read paper · arxiv.org → Models Analysis Jul 12, 2026
Jul 12HF Daily Papers

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning

We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for multimodal reasoning. For each query, the model proposes an answer using the same weights, and issues…

Read paper · arxiv.org → Multimodal System Jul 12, 2026
Jul 12HF Daily Papers

A Vocabulary for Multi-Agent Automated Research Systems

We introduce a vocabulary for automated research systems built from one or more agents to make their design choices easier to describe and compare. The vocabulary specifies 1) who the agents are, 2) what operations are available in the…

Read paper · arxiv.org → Infra System Jul 12, 2026
Jul 11

Quick Automate's case management

AWS wrapped agent runs in first-class "cases" that move through a real lifecycle (Ready, In Progress, Successful, Failed, Pending Resolution), with a creator/processor split so one automation ingests work and many process it in parallel. That's less a research finding than a naming exercise, the one enterprises actually need before they'll trust a fleet of agents with production work and a paper trail.

Read paper · aws.amazon.com → Infra System Jul 11, 2026
Jul 11

UniClawBench checks whether agents can do things, not recite them.

400 bilingual tasks run live in Docker containers, scored step by step across five capabilities (skill use, exploration, long-context reasoning, multimodal understanding, cross-platform coordination), instead of graded off a static answer key an agent can game. Grading criteria stay hidden during the run. Closer to how you'd actually judge an assistant.

Read paper · arxiv.org → Multimodal System Jul 11, 2026
Jul 11

OpenCoF swaps chain-of-thought text for chain-of-frame video.

The model reasons by generating a sequence of images instead of writing out steps, trained on a newly built dataset for the task. The resulting model, Wan-CoF, beat its baseline across four reasoning benchmarks.

Read paper · arxiv.org → Multimodal Dataset Jul 11, 2026
Jul 11HF Daily Papers

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We…

Read paper · arxiv.org → Robotics System Jul 11, 2026
Jul 11HF Daily Papers

ABot-N1: Toward a General Visual Language Navigation Foundation Model

Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolithic policies that…

Read paper · arxiv.org → Robotics Method Jul 11, 2026
Jul 11HF Daily Papers

Towards Autonomous and Auditable Medical Imaging Model Development

Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback. Translating this capability to medical imaging remains difficult…

Read paper · arxiv.org → Science System Jul 11, 2026
Jul 11HF Daily Papers

Predictive Divergence Masks for LLM RL

Reinforcement learning for large language models (LLMs) typically relies on trust-region masks to stabilize off-policy updates. The dominant PPO-style approach uses the sampled-token importance ratio for two criteria: a proximity…

Read paper · arxiv.org → Models Method Jul 11, 2026
Jul 10

Proactive memory agent splits remembering from acting: a separate agent watches a long trajectory and decides what's worth resurfacing and when, instead of hoping the agent doing the work notices before it's buried in context or pushed out entirely.

For anyone building multi-step coding or browsing agents, this is the wall you actually hit around step 50, not model quality.

Read paper · arxiv.org → Agents Method Jul 10, 2026
Jul 10

WebSwarm drops the single-agent, single-trajectory approach to deep research.

It recursively spins up sub-agent teams to cover a search space in parallel, then folds the results back up the tree. A lone ReAct loop runs out of context and patience on genuinely deep-and-wide queries. A swarm doesn't.

Read paper · arxiv.org → Agents Method Jul 10, 2026
Jul 10

LLM judge swap

Researchers held candidate responses fixed across four judgment datasets and only changed which model graded them. The scores moved anyway, no answers touched. If your eval pipeline shows a quality jump after a model upgrade, check whether you upgraded the judge too before you believe it.

Read paper · arxiv.org → Evals Dataset Jul 10, 2026
Jul 10HF Daily Papers

CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation

Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet most systems give the user little control over how a garment should be worn -- its size (loose or fitted), style (e.g.,…

Read paper · arxiv.org → Robotics System Jul 10, 2026
Jul 10HF Daily Papers

4D Human-Scene Reconstruction from Low-Overlap Captures

Existing volumetric capture of dynamic human performance achieves high fidelity with dense camera arrays. However, in real-world scenarios, only a handful of low-overlap cameras are available, which degrades the output quality and leaves…

Read paper · arxiv.org → Evals Method Jul 10, 2026
Jul 10HF Daily Papers

SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding

Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as length, layout…

Read paper · arxiv.org → Robotics Benchmark Jul 10, 2026
Jul 10HF Daily Papers

GRASP: GRanularity-Aware Search Policy for Agentic RAG

Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains challenging for models to decide when…

Read paper · arxiv.org → Evals Benchmark Jul 10, 2026
Jul 10HF Daily Papers

GigaChat Audio: Time-aware Large Audio Language Model

Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time…

Read paper · arxiv.org → Multimodal Method Jul 10, 2026
Jul 10HF Daily Papers

GigaAM Multilingual: Foundation Model for Underrepresented Languages

Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffering from severe data scarcity. This work addresses the challenge of building robust foundation models for underrepresented…

Read paper · arxiv.org → Models Method Jul 10, 2026
Jul 10HF Daily Papers

Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization

Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However, RL algorithms with PPO-Clip are inherently limited by exploration collapse. Subsequent works remain primarily heuristic and…

Read paper · arxiv.org → Evals Method Jul 10, 2026
Jul 9

NVIDIA + LangChain's Deep Agents

Nemotron 3 Ultra, tuned through LangChain's Deep Agents harness, now matches the top closed models on LangChain's own agentic benchmark at roughly a tenth of the inference cost. Nobody retrained the model. The gains came entirely from rebuilding the harness around it: system prompts, tool descriptions, middleware.

Read paper · blogs.nvidia.com → Evals Benchmark Jul 9, 2026
Jul 9

OpenAI re-audits SWE-Bench Pro

OpenAI ran its own audit of the coding benchmark half the industry cites and found a sizable chunk of its tasks are broken: hidden requirements, contradictory instructions, tests too strict for even a correct fix to pass. It's retracting its own recommendation of the eval. The audit process looks solid, but clock who's doing the debunking: a lab that sits near the top of a lot of other coding leaderboards.

Read paper · openai.com → Evals Benchmark Jul 9, 2026
Jul 9

Mistral's Robostral Navigate

An 8B model that takes one RGB camera feed and a plain-language instruction and gets a robot where you told it to go. No LIDAR, no depth sensor, state of the art on the R2R-CE navigation benchmark. Small and single-camera by design. That's the actual news: navigation stacks that used to need a full sensor suite might not for much longer.

Read paper · x.com → Robotics Benchmark Jul 9, 2026
Jul 9

Phone Segmentation and Recognition through Phonological Activation Mapping

Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and…

Read paper · arxiv.org → Multimodal Method Jul 9, 2026
Jul 9

PanoWorld: Real-World Panoramic Generation

In this work, we aim to address the challenge of long-range memory in panoramic world models by exploiting the rotation-equivariant property of omnidirectional representations, where rotation can be treated as an implicit geometric…

Read paper · arxiv.org → Evals Method Jul 9, 2026
Jul 9

Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning

Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to use them for downstream reasoning tasks. We formalize this failure as the \textbf{Knowing--Using Gap}, characterized by…

Read paper · arxiv.org → Models Method Jul 9, 2026
Jul 9

Self-Guided Test-Time Training for Long-Context LLMs

Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does not guarantee effective utilization of long inputs. As input length grows, accuracy often degrades,…

Read paper · arxiv.org → Models Method Jul 9, 2026
Jul 9

Video Generation Models are General-Purpose Vision Learners

Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we…

Read paper · arxiv.org → Multimodal Method Jul 9, 2026
Jul 9

Scalable Visual Pretraining for Language Intelligence

The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and…

Read paper · arxiv.org → Multimodal Method Jul 9, 2026
Jul 9HF Daily Papers

On Locality and Length Generalization in Visual Reasoning

A striking feature of the human visual system is that it ingests visual information through a series of local foveated glimpses, rather than a single global computation. This makes human vision distinctly different from most popular…

Read paper · arxiv.org → Multimodal System Jul 9, 2026
Jul 9HF Daily Papers

OpenLongTail: Generative Scaling of Long-Tail Driving Data

Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected…

Read paper · arxiv.org → Models Dataset Jul 9, 2026
Jul 9HF Daily Papers

DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs

Retired GPUs can form low-cost clusters for LLM inference, but their economic and environmental viability depends heavily on local electricity prices and carbon intensity.

Read paper · arxiv.org → Infra Method Jul 9, 2026
Jul 8

Doomed from the Start

Multi-step LLM agents routinely lock into a trajectory that's going to fail long before the failure is visible, and burn real inference budget getting there anyway. This paper finds the tell shows up early in the agent's internal representations, and a lightweight probe cascade catches it in time to abort. Run agents at any volume and that's compute you're currently spending on episodes that were never going to work.

Read paper · arxiv.org → Agents Method Jul 8, 2026
Jul 8

Your benchmark table lied to you

Most LLM data-analysis benchmarks test fact retrieval on small, clean tables. This paper builds one closer to what an analyst actually opens: large multi-table joins with half-documented columns, plus questions that need outside context to answer. Scores drop hard. A model's spreadsheet-benchmark score isn't a great predictor of how it'll do on your actual warehouse.

Read paper · arxiv.org → Evals Benchmark Jul 8, 2026
Jul 8

Compliance mapping, automated

Matching cloud security controls to the technical metrics that prove you meet them is still mostly manual work. Researchers fine-tuned sentence transformers on 3,499 semantic pairs pulled from five European security standards, and the domain-adapted model does the matching on its own. Narrow use case, but it's exactly the grunt work a compliance team would happily hand off.

Read paper · arxiv.org → Robotics Method Jul 8, 2026
Jul 8

DrugGen 2: A disease-aware language model for enhancing drug discovery

DrugGen-2 generates small molecules conditioned on disease ontology and target protein sequences through fine-tuning GPT-2 with supervised learning and reinforcement learning using GRPO, achieving superior molecular diversity and binding…

Read paper · arxiv.org → Science Method Jul 8, 2026
Jul 8

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models

LongE2V enables high-quality video recovery from sparse event streams by leveraging pre-trained video diffusion priors and addressing temporal stability and frame interpolation challenges.

Read paper · arxiv.org → Multimodal Method Jul 8, 2026
Jul 8HF Daily Papers

Enhancing In-context Panoramic Generation via Geometric-aware Pretraining

Canvas360 is a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with fine-tuning, featuring a large-scale dataset and novel modeling techniques for improved geometric consistency and global…

Read paper · arxiv.org → Evals Dataset Jul 8, 2026
Jul 8HF Daily Papers

A Quantized Native Runtime for On-Device Semantic Audio Generation

A dependency-free runtime enables efficient text-to-music generation on embedded devices through quantization and activation steering while maintaining audio quality.

Read paper · arxiv.org → Multimodal System Jul 8, 2026
Jul 8HF Daily Papers

CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

CausalDS is a benchmark for evaluating causal reasoning in data-science workflows that combines synthetic causal structures with realistic observational data and natural-language stories across Pearl's three rungs of causal inference.

Read paper · arxiv.org → Science Benchmark Jul 8, 2026
Jul 8HF Daily Papers

ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation

ARDY is a streaming generation framework that enables real-time, high-fidelity 3D human motion generation with text and kinematic constraint control through a hybrid representation and two-stage autoregressive transformer denoiser.

Read paper · arxiv.org → Robotics System Jul 8, 2026
Jul 8HF Daily Papers

Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation

A benchmark for scientific lineage reasoning and idea generation is introduced, organizing scientific works as genetic-like Idea Genome objects and evaluating both reasoning and generation capabilities.

Read paper · arxiv.org → Evals Benchmark Jul 8, 2026
Jul 8HF Daily Papers

OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

OPSD-V enhances few-step autoregressive video diffusion models by using real long-video data for temporal context during training, providing dense trajectory-level supervision that improves visual quality and motion dynamics without…

Read paper · arxiv.org → Multimodal Method Jul 8, 2026
Jul 8

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation

Modern Video Object Segmentation (VOS) involves tracking and segmenting user-specified targets. While recent approaches have achieved remarkable performance in single-target scenarios, extending them to multi-target settings typically…

Read paper · arxiv.org → Multimodal Method Jul 8, 2026
Jul 8HF Daily Papers

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This…

Read paper · arxiv.org → Evals Benchmark Jul 8, 2026
Jul 8HF Daily Papers

From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

Pretrained diffusion transformers can be adapted for dense prediction tasks by mapping tokens to task-native outputs instead of generating RGB images, achieving state-of-the-art results with minimal additional parameters.

Read paper · arxiv.org → Multimodal Method Jul 8, 2026
Jul 8HF Daily Papers

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs. These examples suggest that existing…

Read paper · arxiv.org → Multimodal Benchmark Jul 8, 2026
Jul 8HF Daily Papers

A Theory of Contrastive Learning with Natural Images

Why does contrastive learning with simple images and augmentations yield useful representations for downstream tasks? We address this question by analytically computing the optimal representation in terms of a contrastive loss for a range…

Read paper · arxiv.org → Multimodal Method Jul 8, 2026
Jul 8HF Daily Papers

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a forecast. We ask whether internal representations offer a…

Read paper · arxiv.org → Models Method Jul 8, 2026
Jul 8HF Daily Papers

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tailed: new characters, trending entities, post-cutoff events, and more. This…

Read paper · arxiv.org → Multimodal Method Jul 8, 2026
Jul 8HF Daily Papers

MuScriptor: An Open Model for Multi-Instrument Music Transcription

Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic training data, the resulting models generalize poorly,…

Read paper · arxiv.org → Models Method Jul 8, 2026
Jul 7

LLM-as-a-Verifier

A new framework argues that verification, not more pretraining or longer chains of thought, is the next scaling axis: train a model to grade solutions well, and you get a cheaper lever than training it to produce them. Holds up in-paper. Whether it holds outside the benchmarks the authors picked is the open question.

Read paper · arxiv.org → Evals Benchmark Jul 7, 2026
Jul 7

Multiplayer world models

Most "world models" treat every other agent as scenery baked into the environment. This one is the first to condition on multiple agents' action streams directly, so it can attribute a physics outcome to who actually did what. Early-stage, narrow domain, but it's the first real step toward simulators for multi-agent robotics and games instead of solo demos.

Read paper · arxiv.org → Science Method Jul 7, 2026
Jul 7

How Much is Left?

Turns out LLMs linearly encode how many tokens they have left to generate, in the hidden state, before a single word of the answer exists. That's a tidy explanation for why models ramble or truncate in oddly consistent patterns, and a real hook for anyone trying to build length control that doesn't rely on cutting text off after the fact.

Read paper · arxiv.org → Robotics Method Jul 7, 2026
Jul 7

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

SciReasoner is a multimodal scientific foundation model that enables interpretable structural reasoning across proteins, molecules, and crystals by discretizing structural elements into a unified vocabulary for enhanced prediction and…

Read paper · arxiv.org → Science Method Jul 7, 2026
Jul 7

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

LaMem-VLA introduces a latent-memory-native framework that integrates historical experience into vision-language-action reasoning through coordinated memory components operating in the same latent space.

Read paper · arxiv.org → Robotics System Jul 7, 2026
Jul 7

Infinite Worlds with Versatile Interactions

An advanced world modeling system with extended interaction capabilities, real-time processing, diverse interactive elements, and multi-agent behavior control for collaborative virtual environments.

Read paper · arxiv.org → Robotics System Jul 7, 2026
Jul 7

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

LingBot-Video presents a DiT-based video pretraining framework with Mixture-of-Experts architecture, specialized data augmentation, and multi-dimensional reward system for embodied intelligence applications.

Read paper · arxiv.org → Robotics System Jul 7, 2026
Jul 7

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Asynchronous reinforcement learning with single-rollout optimization addresses stability issues in LLM training for complex tasks, outperforming existing methods in coding and reasoning benchmarks.

Read paper · arxiv.org → Evals Benchmark Jul 7, 2026
Jul 7

Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity

Sparse Delta Memory extends gated linear RNNs with sparse addressing to dramatically increase hidden state capacity for improved long-context learning and retrieval while maintaining computational efficiency.

Read paper · arxiv.org → Evals Benchmark Jul 7, 2026
Jul 7HF Daily Papers

Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing

A comparative analysis of softmax attention and recurrent linear-attention architectures examines their expressivity, memory management, and training efficiency across different parameter scales and sequence lengths.

Read paper · arxiv.org → Infra System Jul 7, 2026
Jul 7HF Daily Papers

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma

Reinforcement learning frameworks for large language models face exploration-stability trade-offs, which are addressed through a novel universal objective called Unbounded Positive Asymmetric Optimization that enables stable training with…

Read paper · arxiv.org → Evals System Jul 7, 2026
Jul 7HF Daily Papers

Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE

A novel zero-shot method called Jet-Long enables efficient long-context processing for large language models by dynamically adapting rescaling factors and utilizing a bifocal attention mechanism that maintains high performance across…

Read paper · arxiv.org → Models Method Jul 7, 2026
Jul 7HF Daily Papers

A Sparse and Truncated State Vector Simulator for Peaked Circuits

Peaked quantum circuits can be efficiently simulated classically using sparse state vector representations with vectorized operations and hardware acceleration.

Read paper · arxiv.org → Science Method Jul 7, 2026
Jul 7HF Daily Papers

Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulation

Flow-ERD is a multi-agent traffic simulator that combines agent-type aware flow matching with entropy-regularized distillation to achieve both realistic and diverse motion patterns.

Read paper · arxiv.org → Agents Method Jul 7, 2026
Jul 7HF Daily Papers

KronQ: LLM Quantization via Kronecker-Factored Hessian

Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without retraining. Existing second-order PTQ methods, including GPTQ, construct quantization objectives exclusively from input…

Read paper · arxiv.org → Infra Method Jul 7, 2026
Jul 7HF Daily Papers

Weak-to-Strong Generalization via Direct On-Policy Distillation

Direct On-Policy Distillation transfers reinforcement learning improvements from smaller to larger models by using the policy shift induced by RL as an implicit reward signal, enabling efficient scaling of training without re-running…

Read paper · arxiv.org → Models Method Jul 7, 2026
Jul 7HF Daily Papers

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet the development of multimodal foundation models is constrained by limited access to large-scale, high-quality clinical data.…

Read paper · arxiv.org → Science System Jul 7, 2026
Jul 7HF Daily Papers

From Noisy Traces to Root Causes: Structural Trajectory Analysis and Causal Extraction for Agent Optimization

The optimization of long-horizon agents increasingly relies on reflection-based mechanisms, where a large language model (LLM) acts as an optimizer to diagnose agent failures and improve agent policies. However, real execution traces are…

Read paper · arxiv.org → Agents Analysis Jul 7, 2026
Jul 7HF Daily Papers

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: the same image may be allowed in one product, restricted in…

Read paper · arxiv.org → Multimodal Benchmark Jul 7, 2026
Jul 7HF Daily Papers

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even explicitly…

Read paper · arxiv.org → Evals Benchmark Jul 7, 2026
Jul 7HF Daily Papers

Length Penalties Make Chain-of-Thought Less Monitorable

Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer. In our experiments, training with length penalties does not stop misleading hints from steering…

Read paper · arxiv.org → Models Method Jul 7, 2026
Jul 7HF Daily Papers

SPEAR: A Simulator for Photorealistic Embodied AI Research

Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmability, and rendering speed. We address…

Read paper · arxiv.org → Robotics Method Jul 7, 2026
Jul 7HF Daily Papers

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward reinforcement learning provides weak supervision for…

Read paper · arxiv.org → Multimodal Method Jul 7, 2026
Jul 6

Ditch Adam for interatomic potentials

A new study training NequIP and Allegro models found that matrix-structured optimizers SOAP and SOAP-Muon beat plain Adam on both convergence speed and final accuracy, with the gap widening when force labels are sparse. Optimizer choice has been an afterthought in molecular-simulation AI. This says it shouldn't be.

Read paper · arxiv.org → Models Analysis Jul 6, 2026
Jul 6

CLIP's typography blind spot, patched with no retraining

Write "banana" on a photo of an apple and a CLIP-based model will often agree with the text over its own eyes. Researchers traced the failure to specific attention heads that favor lexical signal over visual signal, then fixed it by intervening on just those heads. Cheap insurance if you're shipping a CLIP classifier anyone could scribble on.

Read paper · arxiv.org → Multimodal Method Jul 6, 2026
Jul 6

From Foundation to Application: Improving VLA Models in Practice

LingBot-VLA 2.0 enhances generalization across tasks and embodiments through expanded data preprocessing and training on diverse robot configurations, extends action space to include whole-body degrees of freedom for complex manipulation…

Read paper · arxiv.org → Robotics Method Jul 6, 2026
Jul 6

Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator

Image2Sim enables scalable embodied navigation training by creating high-fidelity interactive environments from RGB-D images through decoupled 3D spatial anchoring and photorealistic rendering techniques.

Read paper · arxiv.org → Robotics Method Jul 6, 2026
Jul 6

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

PluraMath extends the PolyMath dataset to 18 underrepresented languages, revealing persistent gaps in multilingual mathematical reasoning performance between high-resource and low-resource languages.

Read paper · arxiv.org → Science Dataset Jul 6, 2026
Jul 6

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

Turn-level budgeting strategy for efficient on-policy distillation in long-horizon agent training addresses inefficiencies in full-horizon rollouts and shallow token concentration.

Read paper · arxiv.org → Agents Method Jul 6, 2026
Jul 6HF Daily Papers

Vision as Unified Multimodal Generation

A unified multimodal model formulates computer vision tasks as generation problems using natural language and visual prompts, achieving performance comparable to specialized systems across diverse vision tasks.

Read paper · arxiv.org → Multimodal System Jul 6, 2026
Jul 6HF Daily Papers

Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval Models

MaxSim similarity can exactly replicate inner products between sparse vectors and supports logical operations, with Signed MaxSim extending this capability to real-valued vectors while improving retrieval performance on complex query types.

Read paper · arxiv.org → Evals Benchmark Jul 6, 2026
Jul 6HF Daily Papers

SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models

SIEVE is a structure-aware data selection method for vision-language-action imitation learning that identifies reusable visuo-motor primitives and transition interfaces to improve policy learning efficiency.

Read paper · arxiv.org → Multimodal Method Jul 6, 2026
Jul 6HF Daily Papers

AlayaWorld: Long-Horizon and Playable Video World Generation

AlayaWorld is an open-source framework for creating interactive generative worlds that enables real-time user interaction and supports diverse actions within a modular architecture.

Read paper · arxiv.org → Multimodal System Jul 6, 2026
Jul 6HF Daily Papers

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding

Nemotron-Labs-Diffusion is a tri-mode language model that combines autoregressive, diffusion, and self-speculation decoding to achieve superior throughput and efficiency compared to existing models.

Read paper · arxiv.org → Infra Method Jul 6, 2026
Jul 6HF Daily Papers

RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation

Digital teleoperation replaces physical robot interaction with generative world models to create diverse training data for robotics, enabling efficient zero-shot Sim2Real transfer and improved real-world performance.

Read paper · arxiv.org → Robotics Method Jul 6, 2026
Jul 6HF Daily Papers

RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation

A multi-modal 4D world model generates synchronized RGB, depth, and optical flow data from single RGB-D images and language instructions, enabling efficient robotic manipulation through unified diffusion processes and inverse dynamics…

Read paper · arxiv.org → Robotics Method Jul 6, 2026
Jul 6

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

RoboDojo presents a unified sim-and-real benchmark for evaluating generalist robot manipulation policies across diverse tasks and evaluation dimensions.

Read paper · arxiv.org → Robotics Benchmark Jul 6, 2026
Jul 6

WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence

WildCity presents a large-scale multimodal dataset for urban navigation and spatial representation, enabling research into AI systems that can perceive and reason about city-scale environments similar to human cognitive capabilities.

Read paper · arxiv.org → Robotics Dataset Jul 6, 2026
Jul 6

SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review

Agentic code review framework enhances AI-generated pull requests through iterative review and revision cycles, improving both code quality and issue resolution capabilities.

Read paper · arxiv.org → Multimodal Survey Jul 6, 2026
Jul 6

Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model Failure

World models exhibit long-horizon failures due to kinematic rather than dynamic imagination, as demonstrated by measuring imagined kinematic-consistency error which remains flat while policy rewards collapse across friction boundaries.

Read paper · arxiv.org → Models Analysis Jul 6, 2026
Jul 6HF Daily Papers

PhyMRI-SR: Toward Physics-Aware MRI Image Super-Resolution

MR super-resolution is reformulated as a physics-aware reconstruction problem that dynamically adapts resolution-SNR configurations using Gaussian splatting with prior-aware representations and physics-constrained modeling.

Read paper · arxiv.org → Science Method Jul 6, 2026
Jul 6HF Daily Papers

RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures

RoboTALES introduces a two-stage framework that combines LLM-based planning and VLM-based criticism to improve task-aligned video generation and robotic policy training.

Read paper · arxiv.org → Robotics System Jul 6, 2026
Jul 6HF Daily Papers

Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification

A token-centric dual-view learning framework unifies prompt-based adaptation and cross-view fusion in a frozen vision transformer to improve breast cancer classification from mammography images.

Read paper · arxiv.org → Multimodal System Jul 6, 2026
Jul 6HF Daily Papers

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience the entire…

Read paper · arxiv.org → Evals Survey Jul 6, 2026
Jul 6HF Daily Papers

VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery

Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact exploration. However, in cultural heritage domains such as ancient Greek pottery,…

Read paper · arxiv.org → Multimodal Method Jul 6, 2026
Jul 6HF Daily Papers

UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation

Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on complex prompts that impose substantial demands on users and offer limited expressivity for page…

Read paper · arxiv.org → Multimodal Benchmark Jul 6, 2026
Jul 5

ReContext solves a problem worth naming: your model can hold 128K tokens and still ignore the answer sitting right there in it.

The fix is training-free. It replays the model's own attention signals as a query-conditioned evidence pool before final generation, tested on eight long-context benchmarks across Qwen3-4B, Qwen3-8B, and Llama3-8B. Best average rank on all three, no fine-tuning required.

Read paper · arxiv.org → Evals Benchmark Jul 5, 2026
Jul 5

DemoPSD tackles self-distillation's dirty secret: when a model teaches itself to reason, the "teacher" version often leaks answer-dependent shortcuts the student can't use at test time, collapsing exploration.

The fix blends teacher and student token distributions adaptively instead of forcing dense supervision, then beats GRPO and SDPO on SciKnowEval and holds up better on out-of-distribution GPQA.

Read paper · arxiv.org → Multimodal Benchmark Jul 5, 2026
Jul 5

DramaSR-LRM is the odd one out, and worth it anyway.

Speaker attribution in TV dialogue breaks voice-recognition models on short lines, so this ICML 2026 paper swaps acoustic biometrics for a reasoning model that pulls in visual and linguistic context instead. Comes with a 532K-line, 900-character benchmark dataset (DramaSR-532K) the authors are releasing.

Read paper · arxiv.org → Multimodal Dataset Jul 5, 2026
Jul 5

KVpop -- Key-Value Cache Compression with Predictive Online Pruning

KVpop learns optimal key-value cache eviction by directly supervising keep-or-drop decisions using future-attention targets, achieving high performance with reduced memory usage.

Read paper · arxiv.org → Agents Method Jul 5, 2026
Jul 5

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

A video diffusion framework generates long, multi-view consistent videos by combining temporal and view-wise autoregression through 4D geometric bridging and spatio-temporal distillation techniques.

Read paper · arxiv.org → Multimodal System Jul 5, 2026
Jul 5

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval

Object-aware token merging framework SaMer compresses image-side tokens while preserving query-selectable visual evidence, achieving significant storage reduction and improved retrieval performance.

Read paper · arxiv.org → Multimodal Benchmark Jul 5, 2026
Jul 5

InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization

InternVLA-A1.5 integrates pretrained vision-language models with future prediction in latent space to enable efficient robot manipulation with preserved semantics and long-horizon execution.

Read paper · arxiv.org → Robotics Method Jul 5, 2026
Jul 5

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Analysis of 38,000 hours of real-world agent interactions reveals log-sigmoid scaling laws for performance and exponential learning speed improvements across 134 diverse tasks.

Read paper · arxiv.org → Agents Analysis Jul 5, 2026
Jul 5HF Daily Papers

Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models

A large-scale visuotactile dataset called Deform360 is introduced to study deformable object dynamics, enabling comparison between 2D video and 3D particle world models for robotic manipulation tasks.

Read paper · arxiv.org → Robotics Dataset Jul 5, 2026
Jul 5HF Daily Papers

PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

PixWorld presents a unified pixel-space diffusion approach for 3D reconstruction and generation that overcomes limitations of latent-space methods through direct image-level supervision and geometry-aware feature alignment.

Read paper · arxiv.org → Multimodal Method Jul 5, 2026
Jul 5HF Daily Papers

Vision Pretraining for Dense Spatial Perception

Boundary modeling enables dense spatial perception by learning sub-pixel representations that enhance depth estimation and support embodied AI applications.

Read paper · arxiv.org → Robotics Method Jul 5, 2026
Jul 5HF Daily Papers

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration

A large-scale multimodal tool-use dataset and agent are presented for complex image creation workflows that orchestrate multiple visual tools through multi-turn interactions and hybrid reward optimization.

Read paper · arxiv.org → Multimodal Dataset Jul 5, 2026
Jul 5HF Daily Papers

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

DSpark enhances LLM inference speed by combining parallel draft generation with adaptive verification that reduces waste and improves throughput in high-concurrency settings.

Read paper · arxiv.org → Models Method Jul 5, 2026
Jul 5HF Daily Papers

TREK: Distill to Explore, Reinforce to Refine

TREK expands exploration support for policy optimization by using distillation for exploration rather than imitation, improving performance on challenging mathematical reasoning and agentic tasks.

Read paper · arxiv.org → Science Method Jul 5, 2026
Jul 5HF Daily Papers

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory

Light-Omni is a multimodal agent framework that enables efficient video understanding through dual contextual states, achieving faster and more accurate video processing by eliminating iterative reasoning while maintaining semantic…

Read paper · arxiv.org → Multimodal System Jul 5, 2026
Jul 5HF Daily Papers

GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks

Graph-as-Policy system combines modular robot skills with multi-agent coding to improve reliability in variable automation tasks through parallel simulation refinement.

Read paper · arxiv.org → Robotics System Jul 5, 2026
Jul 5

Teaching LLMs a Low-Resource Language: Enhancing Code Completion in Pharo

Large language models can be adapted for low-resource programming languages through specialized training pipelines and benchmarks, achieving superior code completion performance compared to general-purpose models.

Read paper · arxiv.org → Evals Benchmark Jul 5, 2026
Jul 5

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

HunyuanOCR-1.5 is a lightweight end-to-end vision-language model that enhances OCR capabilities through improved efficiency via DFlash and enhanced capability via Agentic Data Flow, achieving fast inference and broad task coverage.

Read paper · arxiv.org → Multimodal Method Jul 5, 2026
Jul 5

Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES

Byte-pair encoding and Unigram-LM create distinctly different subword vocabularies in chemical language models, with no convergence between the two approaches across diverse corpus types and vocabulary sizes.

Read paper · arxiv.org → Science Dataset Jul 5, 2026
Jul 5HF Daily Papers

PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection

We introduce PAST-TIDE, our stance detection system addressing both subtasks of the StanceNakba Shared Task at NakbaNLP@LREC-COLING 2026. The main idea is statement tuning. We redefine stance as cloze-style masked language modeling (MLM),…

Read paper · arxiv.org → Evals System Jul 5, 2026
Jul 5HF Daily Papers

Trust Region Policy Distillation

Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable…

Read paper · arxiv.org → Models Method Jul 5, 2026
Jul 5HF Daily Papers

Registers Matter for Pixel-Space Diffusion Transformers

Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a problem effectively mitigated by register tokens. As diffusion models increasingly adopt transformer architectures and move…

Read paper · arxiv.org → Multimodal System Jul 5, 2026
Jul 4

Online safety monitoring

A new paper tests the dumbest possible real-time guardrail: take a verifier model's confidence signal, threshold it, calibrate the threshold with risk-control math, raise an alarm when it crosses. On math-reasoning and red-teaming benchmarks, that simple setup matched fancier sequential-hypothesis-testing monitors.

Read paper · arxiv.org → Science Benchmark Jul 4, 2026
Jul 4

LACUNA

Researchers built a testbed that injects fake PII into specific weights of 1B and 7B OLMo models, then checks whether unlearning actually erased it or just buried it. Most SOTA unlearning methods pass every output-level test and still resurface the data under the right probe. When the fact's location in the weights is pinned down precisely, though, a plain gradient-based method erases it cleanly.

Read paper · arxiv.org → Models Method Jul 4, 2026
Jul 4

Program-as-Weights

Compile a fuzzy spec (alert on important log lines, fix malformed JSON, rank by intent) into a small adapter for a frozen local model instead of paying an API per call. The paper's number: a 0.6B Qwen3 running the compiled program matches Qwen3-32B's direct-prompting quality at roughly a fiftieth of the memory, at 30 tokens/sec on a MacBook M3.

Read paper · arxiv.org → Agents Method Jul 4, 2026
Jul 4HF Daily Papers

AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memes

A multimodal sexism identification system for memes uses hierarchical conditional soft-label prediction with vision-language embeddings and a lightweight Gated MLP trained via KL divergence and uncertainty weighting.

Read paper · arxiv.org → Multimodal System Jul 4, 2026
Jul 4HF Daily Papers

dOPSD: On-Policy Self-Distillation for Diffusion Language Models

Diffusion large language models face challenges in reasoning enhancement through post-training, but a novel on-policy self-distillation method using internal denoising trajectories improves mathematical reasoning and code generation…

Read paper · arxiv.org → Science Method Jul 4, 2026
Jul 4HF Daily Papers

UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning

Uni-GUI dataset and UI-MOPD method enable cross-platform GUI agent training by addressing limited data and platform-specific capability degradation through multi-teacher on-policy distillation.

Read paper · arxiv.org → Agents Dataset Jul 4, 2026
Jul 4HF Daily Papers

Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization

A speaker-disentangled syllabic tokenizer regresses perturbed student representations toward clean teacher targets to improve syllable boundary detection and speech language modeling performance.

Read paper · arxiv.org → Multimodal Method Jul 4, 2026
Jul 4HF Daily Papers

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

ResearchStudio-Reel automates research dissemination by composing specialized skills around a shared paper extractor, generating consistent and editable artifacts including posters, videos, and blogs with hard pass/fail quality gates.

Read paper · arxiv.org → Multimodal Method Jul 4, 2026
Jul 4HF Daily Papers

Wan-Streamer v0.2: Higher Resolution, Same Latency

Wan-Streamer v0.2 enhances audio-visual interaction by increasing visual resolution while maintaining low latency through optimized thinker-performer architecture with multi-GPU parallel processing.

Read paper · arxiv.org → Multimodal System Jul 4, 2026
Jul 4HF Daily Papers

ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes

ResearchStudio-Idea provides a skill suite for effective research ideation that combines literature search, novelty checking, and pattern-guided generation to produce traceable research proposals.

Read paper · arxiv.org → Models Benchmark Jul 4, 2026
Jul 4HF Daily Papers

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL

LLM-as-a-Tutor framework extends LLM role from judge to tutor by dynamically adjusting prompt difficulty through pairwise comparison and constraint addition, improving instruction-following performance in reinforcement learning.

Read paper · arxiv.org → Models System Jul 4, 2026
Jul 4HF Daily Papers

SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control

SceneFrom3D generates 3D outdoor scenes by automatically scheduling views from input geometry and controlling object appearance and geometry adherence through identity images and geometry-adherence parameters.

Read paper · arxiv.org → Robotics Method Jul 4, 2026
Jul 4HF Daily Papers

Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models

Flash-BoN improves text-to-image generation efficiency by using inexpensive draft candidates generated through timestep truncation, layer skipping, and activation proxies, followed by multi-stage verification that outperforms existing…

Read paper · arxiv.org → Multimodal Method Jul 4, 2026
Jul 3

Sabotage that ships one PR at a time

a new benchmark (Iterative VibeCoding) has Claude Sonnet 4.5 play attacker against a GPT-4o monitor, spreading a malicious payload across multiple pull requests and timing the final piece for the PR with the best cover. Gradual attacks slipped past a standard diff monitor 93% of the time.

Read paper · arxiv.org → Evals Benchmark Jul 3, 2026
Jul 3

Agents say one thing on the record, another off it

researchers gave 10 models a public channel and a private "off the record" channel inside multi-agent debates, no instructions to behave differently in either. Baseline divergence between the two was about 3%.

Read paper · arxiv.org → Agents Benchmark Jul 3, 2026
Jul 3

The Polymarket test nobody wants to hear

a small real-money forecasting study found most people paired with an AI either just copy its answer or use it to rubber-stamp a guess they'd already made, and both groups do worse than the model alone. A minority beat the model, and what set them apart wasn't IQ or which model they used. It was perspective-taking, intellectual humility, and curiosity.

Read paper · arxiv.org → Models Analysis Jul 3, 2026
Jul 3HF Daily Papers

MANCE: Manifold Aware Concept Erasure

Manifold constraint hypothesis enables improved concept erasure by projecting updates onto estimated representation manifolds, achieving state-of-the-art results in nonlinear concept removal.

Read paper · arxiv.org → Models Method Jul 3, 2026
Jul 3HF Daily Papers

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

OmniOpt presents a unified framework for optimizer selection in large-scale model training by combining meta-pipeline transformations, norm-constrained linear minimization oracles, and a cross-domain benchmark to systematically analyze…

Read paper · arxiv.org → Evals Survey Jul 3, 2026
Jul 3HF Daily Papers

Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification

Automated safety testing framework Vera uses a three-stage pipeline to identify and test safety risks in LLM agents through structured risk taxonomies, combinatorial case generation, and adaptive sandbox execution with evidence-based…

Read paper · arxiv.org → Safety System Jul 3, 2026
Jul 3HF Daily Papers

CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation

CGGS is a text-to-3D framework that enhances 3D-content-awareness and addresses geometric distortions through a multi-stage approach involving ego-centric generation, layout decoration, and geometric refinement.

Read paper · arxiv.org → Multimodal System Jul 3, 2026
Jul 3HF Daily Papers

Attending to Multimodal Generation One Token at a Time

Multimodal large language models exhibit distinct attention patterns during generation, with attention to visual and textual modalities shifting based on semantic requirements, and these patterns can be leveraged to improve task…

Read paper · arxiv.org → Multimodal Method Jul 3, 2026
Jul 3HF Daily Papers

SiamJEPA: On the Role of Siamese Student Encoders in JEPA

Siamese student encoders in JEPA models improve representation separability and training efficiency through effective regularization, outperforming single-encoder variants and MAE under limited training budgets.

Read paper · arxiv.org → Infra Method Jul 3, 2026
Jul 3HF Daily Papers

Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs

Arabic language models exhibit dialect-specific neural representations that can be manipulated at inference time to control dialectal output without requiring additional training.

Read paper · arxiv.org → Robotics Method Jul 3, 2026
Jul 3HF Daily Papers

CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation

CineMobile enables efficient image-to-video generation on mobile devices through distillation-guided pruning, diffusion distillation, and hybrid quantization techniques while maintaining visual quality and achieving significant speedup.

Read paper · arxiv.org → Multimodal Method Jul 3, 2026
Jul 3HF Daily Papers

TESSERA v2: Scaling Pixel-wise Earth Foundation Models

Large-scale controlled experiments reveal optimal scaling strategies for Earth-observation foundation models, enabling efficient training and deployment through encoder growth, downstream performance selection, and model distillation.

Read paper · arxiv.org → Robotics Method Jul 3, 2026
Jul 3HF Daily Papers

OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

OmniTacTune enables efficient adaptation of tactile feedback to visual robot policies through a two-stage reinforcement learning approach that improves success rates in contact-rich manipulation tasks.

Read paper · arxiv.org → Robotics Method Jul 3, 2026
Jul 2

GeneBench-Pro

OpenAI's new benchmark hands models 129 synthetic biology problems that would take a genomics postdoc 20 to 40 hours to work through: not running the stats, but deciding which analysis a messy dataset can actually support. Top score on the board is GPT-5.6 Sol Pro at 31.5%. Claude Opus 4.8 landed at 16%, Gemini 3.5 Flash at 8.1%. Builder read: scientific judgment is still the gap the benchmark race hasn't closed.

Read paper · openai.com → Science Dataset Jul 2, 2026
Jul 2

Human vs. LLM research taste

A new paper checks where model-generated research ideas land against what actual papers propose, and the same pattern shows up across every LLM tested: models cluster around "bridge" ideas, stitching two known methods together, while human researchers spread across a much wider set of ways to frame a gap. The ideas aren't bad. They're narrow.

Read paper · arxiv.org → Models Method Jul 2, 2026
Jul 2

Coding-agent benchmarks, audited

GSO, SWE-Perf, and SWE-fficiency all score agents on real-repo performance optimization. Replay the reference patches and most don't even hold up: 39 of 102 GSO tasks pass cleanly, 11 of 140 on SWE-Perf, 411 of 498 on SWE-fficiency. Rankings flip depending on scoring rules (eight submissions disagree on 9 of 28 head-to-heads), and SWE-fficiency piles 58.5% to 82.8% of its score weight onto the ten hardest tasks.

Read paper · arxiv.org → Evals Benchmark Jul 2, 2026
Jul 2HF Daily Papers

CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training

A 3D latent diffusion model for chest CT generation that achieves high-fidelity results while enabling control over clinical attributes through metadata conditioning and reinforcement learning refinement.

Read paper · arxiv.org → Robotics Method Jul 2, 2026
Jul 2HF Daily Papers

Perceptual Flow Matching for Few-Step Generative Modeling

Perceptual Flow Matching enables efficient few-step generation by supervising flow matching in perceptual feature space, achieving high-quality results with reduced sampling steps and improved accuracy.

Read paper · arxiv.org → Models Method Jul 2, 2026
Jul 2HF Daily Papers

PixCon: Clean-Positive Contrastive Learning for Foundation-Model Semi-Supervised Segmentation

PixCon is a semi-supervised semantic segmentation framework that uses clean-positive pixel-contrastive learning with per-class memory banks to improve accuracy over existing methods.

Read paper · arxiv.org → Agents System Jul 2, 2026
Jul 2HF Daily Papers

PraMem: Practice-derived Experiential Memory for Long-horizon Behavior Prediction

Long-horizon behavior prediction is enhanced through PraMem, which transforms lengthy historical sequences into experiential memory for improved accuracy.

Read paper · arxiv.org → Agents Method Jul 2, 2026
Jul 2HF Daily Papers

Bibby AI: An Editor-Native Agentic Platform for Academic Research, Writing, and Publishing

Bibby AI is an editor-native platform that consolidates the academic writing workflow into a unified Research-Write-Publish pipeline, streamlining literature management, citation insertion, and formatting through integrated tools and…

Read paper · arxiv.org → Agents System Jul 2, 2026
Jul 2HF Daily Papers

MentalThink: Shaping Thoughts in Mental SVG World

MentalThink enables multimodal large language models to perform visual-symbolic reasoning by generating and interpreting SVG code as an executable intermediate representation for spatial problem-solving.

Read paper · arxiv.org → Multimodal Method Jul 2, 2026
Jul 2HF Daily Papers

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment

A supervised contrastive alignment framework maps WavLM embeddings from English and Mandarin into a shared clinical space for depression detection, addressing cross-lingual generalization challenges and revealing performance artifacts…

Read paper · arxiv.org → Multimodal System Jul 2, 2026
Jul 2HF Daily Papers

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling

Hierarchical Landmark Sparse Attention enables efficient long-context language modeling by learning chunk selection end-to-end, achieving performance comparable to full attention while extrapolating beyond training context lengths.

Read paper · arxiv.org → Models Method Jul 2, 2026
Jul 2HF Daily Papers

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

Flex-Forcing enables video diffusion models to operate under both bidirectional and autoregressive generation regimes through a flexible chunking mechanism over temporal and denoising steps, improving video quality and inference speed.

Read paper · arxiv.org → Multimodal Method Jul 2, 2026
Jul 2HF Daily Papers

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning

A parallelized autoregressive framework for dense video captioning that improves generation efficiency by exploiting weak local dependencies across temporally distinct events while maintaining temporal grounding accuracy.

Read paper · arxiv.org → Multimodal System Jul 2, 2026
Jul 2HF Daily Papers

SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe

A minimal viable pipeline for skill optimization is proposed through Zeroth-Order optimization formalization, eliminating redundancies while maintaining convergence and generalization through trajectory exploration, consensus mining, and…

Read paper · arxiv.org → Agents System Jul 2, 2026
Jul 2HF Daily Papers

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

Large Audio-Language Models exhibit systematic generative biases in realistic scenarios when evaluated through open-ended tasks using human-recorded speech, with bias magnitude varying significantly by task and triggered by gender and…

Read paper · arxiv.org → Multimodal Benchmark Jul 2, 2026
Jul 2HF Daily Papers

Automating the Design of Embodied Agent Architectures

Automated agent architecture search demonstrates potential for improving embodied agent performance while revealing challenges related to optimization signals, local optima, and credit assignment in simulation-based training.

Read paper · arxiv.org → Robotics System Jul 2, 2026
Jul 2HF Daily Papers

Vidu S1: A Real-Time Interactive Video Generation Model

Vidu S1 is a real-time interactive video generation model that supports voice-controlled digital character animation with infinite-length output and high frame rate on consumer hardware.

Read paper · arxiv.org → Robotics Method Jul 2, 2026
Jul 2HF Daily Papers

Spectral Rewiring for Exploration, Purification, and Model Merging

Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter updates create two deployment-relevant bottlenecks: suppressed reasoning performance, often reflected by premature…

Read paper · arxiv.org → Models Method Jul 2, 2026
Jul 1

Metacognitive RL

LLMs are bad at knowing what they don't know: confident hallucinations, no sense of their own knowledge boundary. This paper trains models with RL on metacognitive feedback and gets stated confidence to actually track correctness. Builder read: closer to a model you can trust to abstain instead of bluff.

Read paper · arxiv.org → Evals Method Jul 1, 2026
Jul 1

Table reading errors models parse table structure fine, then still cite the wrong cell.

The error isn't comprehension, it's reference. Matters if you're shipping anything that answers questions over spreadsheets or financial tables: right-looking answer, wrong number. The paper says it's fixable, not fundamental.

Read paper · arxiv.org → Models Method Jul 1, 2026
Jul 1

Skill composition for agents

instead of re-deriving "set up a sandbox" or "run the test suite" from scratch every session, agents assemble reusable skill packages and compose them for new tasks. Same idea Anthropic ships as Skills, now with a paper behind it.

Read paper · arxiv.org → Agents Benchmark Jul 1, 2026
Jul 1

AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition

Research evaluates 30 augmentation strategies for vein recognition across multiple datasets and architectures, revealing inconsistencies between accuracy and adversarial security while highlighting dataset-specific effectiveness variations.

Read paper · arxiv.org → Safety Dataset Jul 1, 2026
Jul 1

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Autonomous agents evaluate policy improvement through iterative editing within fixed budgets, revealing that successful policy evolution requires both task-specific mechanisms and feedback-constrained refinement.

Read paper · arxiv.org → Evals Benchmark Jul 1, 2026
Jul 1

AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents

A bounded contract approach for long-horizon LLM agents uses typed retrieval to assemble fresh prompts, enabling isolated analysis of memory components and demonstrating improved performance in complex decision-making tasks.

Read paper · arxiv.org → Evals Benchmark Jul 1, 2026
Jul 1

From SRA to Self-Flow: Data Augmentation or Self-Supervision?

Research investigates the mechanisms behind self-alignment methods in diffusion transformers, finding that performance improvements stem primarily from data augmentation along the noise dimension rather than token interactions between…

Read paper · arxiv.org → Multimodal Analysis Jul 1, 2026
Jul 1

SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use

SkillCoach is a self-evolving rubric framework that evaluates and improves agentic skill-use by analyzing skill selection, following, composition, and reflection processes, providing better supervision than outcome-only metrics.

Read paper · arxiv.org → Multimodal Benchmark Jul 1, 2026
Jul 1HF Daily Papers

PACE: A Proxy for Agentic Capability Evaluation

PACE is a framework that predicts expensive agentic LLM benchmark performance using a small subset of atomic evaluation instances, achieving high accuracy at a fraction of the cost.

Read paper · arxiv.org → Evals Benchmark Jul 1, 2026
Jul 1HF Daily Papers

Representation Distribution Matching for One-Step Visual Generation

Representation Distribution Matching enables high-quality image generation by matching feature distributions under pretrained encoders, with improved performance through optimized batch sizes and multi-encoder evaluation metrics.

Read paper · arxiv.org → Multimodal Benchmark Jul 1, 2026
Jul 1HF Daily Papers

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

Task-Agnostic Pretraining framework trains robotic models using self-supervised inverse dynamics on unlabeled data followed by lightweight language grounding, achieving superior performance with minimal expert demonstrations.

Read paper · arxiv.org → Robotics System Jul 1, 2026
Jul 1HF Daily Papers

Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling

MrFlow accelerates text-to-image diffusion by combining low-resolution generation with pixel-space super-resolution and noise injection, achieving up to 25x speedup without training or runtime modifications.

Read paper · arxiv.org → Multimodal System Jul 1, 2026
Jul 1HF Daily Papers

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

WorldDirector enables controllable video generation with persistent object memory by decoupling semantic motion planning from visual rendering through LLM coordination of 3D trajectories and camera movements.

Read paper · arxiv.org → Robotics Method Jul 1, 2026
Jul 1HF Daily Papers

Denser neq Better: Limits of On-Policy Self-Distillation for Continual Post-Training

On-policy self-distillation in continual post-training accelerates in-domain specialization but fails to prevent forgetting and can collapse in out-of-distribution scenarios, indicating that on-policy data alone is insufficient for…

Read paper · arxiv.org → Models Method Jul 1, 2026
Jul 1HF Daily Papers

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

Vision-Language Models struggle with domain adaptation in specialized spatio-temporal video grounding tasks, highlighting limitations in zero-shot generalization and in-context learning capabilities.

Read paper · arxiv.org → Multimodal Benchmark Jul 1, 2026
Jul 1HF Daily Papers

Optimizing Visual Generative Models via Distribution-wise Rewards

A novel reinforcement learning framework for visual generation uses distribution-wise rewards to improve image diversity and quality while addressing mode collapse and computational efficiency issues.

Read paper · arxiv.org → Multimodal System Jul 1, 2026
Jul 1HF Daily Papers

AgenticDataBench: A Comprehensive Benchmark for Data Agents

A comprehensive benchmark named AgenticDataBench is introduced to evaluate data agents across diverse domains with fine-grained task annotations and skill-based coverage metrics.

Read paper · arxiv.org → Evals Dataset Jul 1, 2026
Jul 1

WARP: Weight-Space Analysis for Recovering Training Data Portfolios

WARP is a framework that infers training data compositions from released model weights by analyzing geometric footprints in weight space through model merging and feature extraction.

Read paper · arxiv.org → Evals System Jul 1, 2026
Jul 1

Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment

Geo-Anchored Cloud Removal framework addresses semantic drift in cloud removal by combining physically grounded residual inversion with semantic manifold constraints from vision foundation models.

Read paper · arxiv.org → Multimodal System Jul 1, 2026
Jul 1

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers

OrbitQuant enables efficient post-training quantization for diffusion transformers by using a normalized rotated basis that eliminates the need for recalibration across different timesteps and modalities.

Read paper · arxiv.org → Multimodal Method Jul 1, 2026
Jul 1

VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon

VLA-Corrector addresses limitations of action chunking in vision-language-action models by introducing a lightweight latent-space vision monitor that enables adaptive corrective replanning, improving robustness in contact-rich manipulation…

Read paper · arxiv.org → Robotics Method Jul 1, 2026
Jul 1HF Daily Papers

Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

Embodied.cpp is a portable C++ runtime that enables efficient deployment of vision-language-action and world-action models across heterogeneous edge devices through modular execution layers and optimized inference.

Read paper · arxiv.org → Robotics System Jul 1, 2026
Jul 1HF Daily Papers

EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots

EVA-Client is an open-source framework that unifies real-robot policy deployment, data collection, and evaluation through a component-decoupled architecture with inspectable execution workflows.

Read paper · arxiv.org → Robotics Dataset Jul 1, 2026
Jul 1HF Daily Papers

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

World models for robotic policy evaluation are systematically studied through a new benchmark, revealing that long-horizon rollout consistency and robot-specific controllability are more important than short-term visual realism for…

Read paper · arxiv.org → Robotics Benchmark Jul 1, 2026
Jul 1HF Daily Papers

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction

A dual-loop framework named Mastermind is introduced that separates strategy learning from task-specific experience to improve vulnerability reproduction capabilities in software engineering agents.

Read paper · arxiv.org → Agents System Jul 1, 2026
Jul 1HF Daily Papers

Gemma 4 Technical Report

Gemma 4 introduces efficient, multimodal language models with diverse architectures, enhanced reasoning capabilities, and improved performance across various tasks.

Read paper · arxiv.org → Multimodal System Jul 1, 2026
Jul 1HF Daily Papers

PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

A minimalist pixel-space diffusion transformer using plain ViT architecture directly processes 3D point map patches conditioned on image tokens from DINOv3, outperforming complex latent-based models while maintaining simplicity and…

Read paper · arxiv.org → Multimodal System Jul 1, 2026
Jul 1HF Daily Papers

ACID: Action Consistency via Inverse Dynamics for Planning with World Models

ACID is a decision-time planning framework that enhances action-conditioned world models by enforcing cycle action consistency to improve trajectory realism and reduce computational requirements.

Read paper · arxiv.org → Agents System Jul 1, 2026
Jul 1HF Daily Papers

Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training

Reinforcement learning adaptation in transformer models shows highly concentrated improvements in specific middle layers rather than uniform parameter updates across all layers.

Read paper · arxiv.org → Models Method Jul 1, 2026
Jul 1HF Daily Papers

Rank-Then-Act: Reward-Free Control from Frame-Order Progress

Rank-Then-Act framework learns control policies from video demonstrations using a vision-language model as an ordinal scorer with correlation-based rewards, enabling stable cross-task transfer without environment rewards.

Read paper · arxiv.org → Robotics System Jul 1, 2026
Jul 1HF Daily Papers

Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition

RCORE addresses object-driven shortcuts in zero-shot compositional action recognition by using co-occurrence prior regularization and temporal order regularization to improve compositional generalization.

Read paper · arxiv.org → Models Method Jul 1, 2026
Jul 1HF Daily Papers

Video-Oasis: Rethinking Evaluation of Video Understanding

Video-Oasis diagnostics reveal that half of existing video benchmarks can be solved without visual input, exposing significant capability gaps in current video understanding models.

Read paper · arxiv.org → Multimodal Benchmark Jul 1, 2026