Research — June 2026.

The 274 papers the radar tracked in June 2026. The current radar lives on the research page. Back to the radar →

Jun 30

Agents-A1

A 35B MoE model matching trillion-parameter performance by scaling horizon, not parameter count: longer trajectories, broader heterogeneous agent abilities. The claim is architectural. If it reproduces off the paper's own evals, compute allocation for agentic tasks shifts.

Read paper · arxiv.org → Evals Benchmark Jun 30, 2026
Jun 30

WorldEvolver

Agents that self-refine their internal world models to make better predictions before acting. The motivation is sharper than it first sounds: bad foresight isn't just noise. An inaccurate world model can actively degrade decisions, not just add uncertainty. Worth watching whether it holds on harder, real-world planning tasks outside the paper's benchmarks.

Read paper · arxiv.org → Evals Benchmark Jun 30, 2026
Jun 30

Pessimism's Paradox

Conservative offline training amplifies reward hacking during online adaptation. Not suppresses it. The conventional view is that keeping a policy close to well-supported behavior gives you a safer base for online RL. This paper challenges that, both empirically and mechanistically. If you're relying on conservative pretraining as a safety floor before online fine-tuning, read it.

Read paper · arxiv.org → Safety Method Jun 30, 2026
Jun 30HF Daily Papers

Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination

Graph-PRefLexOR, a graph-native reasoning model trained with Group Relative Policy Optimization, improves materials science hypothesis generation through structured phases of mechanism exploration, graph construction, pattern extraction,…

Read paper · arxiv.org → Science Method Jun 30, 2026
Jun 30HF Daily Papers

Cross-Domain Generalization Failure in Lightweight Intrusion Detection Models for IIoT Networks

Lightweight machine learning models for IIoT intrusion detection show limited generalization across networks due to reliance on coarse port-category features and imbalanced class distributions, with adversarial robustness not correlating…

Read paper · arxiv.org → Evals Method Jun 30, 2026
Jun 30HF Daily Papers

Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising

Page-level slide personalization is addressed through a novel framework that formulates the problem as inverse planning and uses a multi-agent reinforcement learning approach to learn design intents without requiring specific tool…

Read paper · arxiv.org → Agents System Jun 30, 2026
Jun 30HF Daily Papers

NoPA: Non-Parametric Online 3D Scene Graph Generation

NoPA introduces a non-parametric distribution-based approach for real-time 3D scene graph generation that preserves geometric details while maintaining computational efficiency through kernel density estimates and particle-based object…

Read paper · arxiv.org → Multimodal Method Jun 30, 2026
Jun 30HF Daily Papers

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning

Asymmetric Mutual Variational Learning addresses train-inference mismatch in multimodal reasoning by using bidirectional calibration to prevent answer leakage and improve latent-space stability.

Read paper · arxiv.org → Multimodal Method Jun 30, 2026
Jun 30HF Daily Papers

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving

ELDR is an expert-locality-aware decode router for prefill-decode disaggregated Mixture-of-Experts serving that improves performance by predicting expert activations and routing requests accordingly.

Read paper · arxiv.org → Infra Method Jun 30, 2026
Jun 30HF Daily Papers

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

Memory plays a crucial role in LLM-based agents, but retrieved memories can cause sycophancy issues where agents over-align with users at the expense of factual accuracy, necessitating new evaluation benchmarks that assess memory's impact…

Read paper · arxiv.org → Evals Benchmark Jun 30, 2026
Jun 30HF Daily Papers

The State-Prediction Separation Hypothesis

Separating state prediction from token prediction in Transformers improves language modeling performance and efficiency across different scales.

Read paper · arxiv.org → Infra Method Jun 30, 2026
Jun 30HF Daily Papers

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

A unified framework named Perceive-to-Reason (P2R) is introduced that separates visual perception from reasoning in vision-language models through a two-stage process, improving fine-grained visual reasoning performance on high-resolution…

Read paper · arxiv.org → Multimodal System Jun 30, 2026
Jun 30HF Daily Papers

CausalMix: Data Mixture as Causal Inference for Language Model Training

CausalMix addresses limitations in LLM data mixing by formulating mixture optimization as a causal inference problem, enabling dynamic adaptation to shifting data distributions without costly retraining.

Read paper · arxiv.org → Models Method Jun 30, 2026
Jun 30HF Daily Papers

Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts

Vision-Language-Action models can be efficiently adapted to new environments using a single demonstration through weight vector arithmetic that isolates domain-specific information via subspace alignment.

Read paper · arxiv.org → Multimodal Method Jun 30, 2026
Jun 30HF Daily Papers

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

ABot-M0.5 is a World Action Model for mobile manipulation that improves performance through temporal granularity alignment, action space disentanglement, and train-test consistency in autoregressive prediction.

Read paper · arxiv.org → Robotics Method Jun 30, 2026
Jun 30HF Daily Papers

Valdi: Value Diffusion World Models

Value Diffusion World Models combine end-to-end online training with latent diffusion dynamics to enable fast, uncertain dynamics prediction for Model Predictive Control in reinforcement learning environments.

Read paper · arxiv.org → Robotics Method Jun 30, 2026
Jun 30HF Daily Papers

Autonomous Scientific Discovery via Iterative Meta-Reflection

An autonomous scientific discovery framework uses large language models and dynamic code generation to conduct open-ended research while maintaining statistical rigor through meta-reflection and multimodal data processing.

Read paper · arxiv.org → Multimodal System Jun 30, 2026
Jun 30HF Daily Papers

VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement

VideoSearch-R1 is an agentic framework that iteratively retrieves videos and refines search queries using continuous latent space refinement and policy optimization for improved video moment retrieval and temporal grounding.

Read paper · arxiv.org → Multimodal Benchmark Jun 30, 2026
Jun 30HF Daily Papers

RepoRescue: An Empirical Study of LLM Agents on Whole-Repository Compatibility Rescue

LLM agents can successfully adapt legacy software repositories to modern environments through compatibility rescue, with collaborative approaches achieving higher success rates than individual systems.

Read paper · arxiv.org → Infra System Jun 30, 2026
Jun 30HF Daily Papers

Discrete Diffusion Language Models for Interactive Radiology Report Drafting

Diffusion language models match or exceed autoregressive models in medical visual question answering while offering faster decoding and bidirectional text editing capabilities.

Read paper · arxiv.org → Science Analysis Jun 30, 2026
Jun 30

Scaling Laws for Grid-Based Approximate Nearest Neighbor Search in High Dimensions

Grid-based multiprobe algorithms demonstrate superior dimensional scaling properties compared to graph-, tree-, and partitioning-based methods for approximate nearest neighbor search, making them competitive for high-dimensional and…

Read paper · arxiv.org → Models Method Jun 30, 2026
Jun 30

AutoMem: Automated Learning of Memory as a Cognitive Skill

Memory management in large language models is treated as a trainable skill through a framework that automates both memory structure optimization and proficiency enhancement, leading to significant performance improvements in long-horizon…

Read paper · arxiv.org → Agents System Jun 30, 2026
Jun 30HF Daily Papers

Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

Logit-Contribution Scoring (LOCOS) identifies attention heads responsible for non-literal context synthesis in large language models by measuring their output-value circuit's contribution to answer tokens, outperforming existing methods on…

Read paper · arxiv.org → Evals Benchmark Jun 30, 2026
Jun 30HF Daily Papers

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

MultAttnAttrib is a training-free multimodal attribution method that locates source evidence in documents using attention heads and calibrated thresholds, achieving superior accuracy and efficiency compared to existing approaches.

Read paper · arxiv.org → Multimodal Method Jun 30, 2026
Jun 30HF Daily Papers

Multi-Turn Agentic Scientific Literature Search via Workflow Induction

paper pilot is a multi-turn literature search agent that uses executable workflows to improve search accuracy and reduce errors by incorporating user feedback and controlled workflow corruption training.

Read paper · arxiv.org → Robotics Method Jun 30, 2026
Jun 30HF Daily Papers

When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers

Research addresses cache management for LLM agents by formalizing semantic cache replacement as an online problem with switching costs, proposing SOLAR—a learning-augmented framework that outperforms traditional heuristics through…

Read paper · arxiv.org → Evals Benchmark Jun 30, 2026
Jun 30HF Daily Papers

RuleChef: Grounding LLM Task Knowledge in Human-Editable Rules

RuleChef utilizes large language models to generate and iteratively improve executable rules for NLP tasks through example-based learning and human feedback, resulting in fast and inspectable rule systems.

Read paper · arxiv.org → Infra System Jun 30, 2026
Jun 30HF Daily Papers

Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs

Splash is a mask-isolated tactile alignment learning framework that enables multimodal LLMs to acquire tactile sensing capabilities without sacrificing vision-language reasoning through selective parameter updating that prevents…

Read paper · arxiv.org → Multimodal System Jun 30, 2026
Jun 30HF Daily Papers

NeuroCogMap Reveals Cognitive Organization of Large Language Models

Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognition. Yet although LLMs exhibit broad cognitive-like…

Read paper · arxiv.org → Infra System Jun 30, 2026
Jun 29

Govern the Repository, Not the Agent

The field evaluates coding agents one at a time on isolated benchmark tasks. This paper argues that's the wrong unit. Agents that each pass their own tests still leave repos accumulating technical debt no per-agent score catches: conflicting assumptions, ownership gaps, drift. The authors propose ecosystem-level risk metrics as the actual measurement frame.

Read paper · arxiv.org → Safety Benchmark Jun 29, 2026
Jun 29

Agent-Native Immune System

Perimeter security and training-time alignment both sit outside the agent's runtime. This paper proposes moving defense inside: a taxonomy and architecture for detecting prompt injection, memory poisoning, and tool-call manipulation as they happen. Early work, but the engineering section is concrete enough to borrow from if you're building agentic infrastructure with real security requirements.

Read paper · arxiv.org → Robotics Survey Jun 29, 2026
Jun 29

Google Paper Assistant

Human peer review can't keep pace with the volume of AI-assisted research. Google benchmarks their Paper Assistant tool against human reviewers, measuring coverage and accuracy on actual submissions. The specific numbers need the full paper; the open question (does automated review hold up in specialized subfields, where human reviewer availability is already thin?) is what makes this worth watching.

Read paper · arxiv.org → Evals Survey Jun 29, 2026
Jun 29

DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation

DataEvolver is a self-evolving multi-agent framework that improves text-rich image generation by leveraging feedback from rejected samples to iteratively enhance data quality.

Read paper · arxiv.org → Multimodal System Jun 29, 2026
Jun 29

MuSViT: A Foundation Vision Model for Sheet Music Representation

MuSViT is a vision transformer-based foundation model pre-trained on millions of sheet music pages that demonstrates superior performance in music score recognition and symbol detection tasks through both linear probing and fine-tuning…

Read paper · arxiv.org → Multimodal Method Jun 29, 2026
Jun 29

GEAR: Guided End-to-End AutoRegression for Image Synthesis

GEAR trains a vector-quantized tokenizer and autoregressive generator jointly end-to-end using representation alignment, overcoming non-differentiability issues through a dual read-out approach that improves convergence speed and feature…

Read paper · arxiv.org → Multimodal Method Jun 29, 2026
Jun 29

Multi-Block Diffusion Language Models

Multi-Block Diffusion Language Models extend single-block diffusion to concurrent block decoding with improved training strategies and optimized decoding algorithms.

Read paper · arxiv.org → Models Method Jun 29, 2026
Jun 29

TerraDiT-Ω: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial Primitive

TerraDiT-Ω generates satellite imagery from native geospatial primitives using Geometry-Aware Local Attention, enabling flexible conditioning and improved downstream geospatial tasks.

Read paper · arxiv.org → Robotics Method Jun 29, 2026
Jun 29

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding

Speculative decoding with adaptive block size selection improves inference efficiency by predicting optimal block sizes from prefilling representations, achieving significant speedup with minimal overhead.

Read paper · arxiv.org → Infra Method Jun 29, 2026
Jun 29HF Daily Papers

Xiaomi-GUI-0 Technical Report

A native multimodal GUI agent trained in real-device environments demonstrates superior performance and stability compared to traditional benchmark-based approaches.

Read paper · arxiv.org → Multimodal Benchmark Jun 29, 2026
Jun 29HF Daily Papers

MemLearner: Learning to Query Context memory for Video World Models

MemLearner improves video world models by using learning-based adaptive context querying with query tokens to enhance scene consistency and memory in long video sequences with occlusions and dynamic objects.

Read paper · arxiv.org → Multimodal Method Jun 29, 2026
Jun 29HF Daily Papers

PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking

Multi-turn visual reasoning agents suffer from entangled reasoning and perception that cause redundant trajectories; PixelEyes addresses this by decoupling these processes through mask-guided search and semantic-region breadth-first…

Read paper · arxiv.org → Multimodal Method Jun 29, 2026
Jun 29HF Daily Papers

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity

Seed2.0 addresses complex real-world tasks by tackling long-tail knowledge and complex instruction following challenges while enhancing reasoning, visual understanding, and search capabilities through a robust evaluation framework grounded…

Read paper · arxiv.org → Multimodal Benchmark Jun 29, 2026
Jun 29HF Daily Papers

ASPIRE: Agentic /Skills Discovery for Robotics

ASPIRE is a continual learning system that autonomously develops and refines robot control programs through iterative exploration, achieving superior performance and zero-shot generalization in manipulation and household tasks while…

Read paper · arxiv.org → Robotics System Jun 29, 2026
Jun 29HF Daily Papers

AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

AutoTrainess enables autonomous language model training by providing structured agent-computer interfaces that guide planning, data preparation, training, evaluation, and logging operations more effectively than traditional command-line…

Read paper · arxiv.org → Evals Benchmark Jun 29, 2026
Jun 29HF Daily Papers

AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation

AtomiMed presents a novel evaluation framework for medical report generation that decomposes clinical narratives into atomic facts and uses an agentic cross-verification process to improve accuracy assessment beyond traditional metrics.

Read paper · arxiv.org → Science Benchmark Jun 29, 2026
Jun 29HF Daily Papers

SpheRoPE: Zero-Shot Optimization-Free 360 Panorama Generation with Spherical RoPE

A novel zero-shot framework injects spherical priors into pre-trained diffusion transformers for 360 panoramic generation, using spherical RoPE and semantic distortion guidance to overcome topological constraints without training or…

Read paper · arxiv.org → Models System Jun 29, 2026
Jun 29HF Daily Papers

TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

TRIAGE introduces a role-typed credit assignment framework that enhances agentic reinforcement learning by providing more nuanced credit assignment than standard GRPO methods.

Read paper · arxiv.org → Agents System Jun 29, 2026
Jun 29HF Daily Papers

InstanceControl: Controllable Complex Image Generation without Instance Labeling

InstanceControl enables multi-instance image generation by using vision-language models to establish instance-level correspondences between text prompts and visual conditions, while employing adaptive mask refinement for improved accuracy.

Read paper · arxiv.org → Robotics Method Jun 29, 2026
Jun 29HF Daily Papers

Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning

A reinforcement learning approach called MRPO is introduced to improve clinical image reasoning by addressing cascading errors through step-wise process rewards, demonstrating superior performance over existing methods.

Read paper · arxiv.org → Science Method Jun 29, 2026
Jun 29HF Daily Papers

Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue

Vision-language models struggle to distinguish between shared and interpreted visual information in dialogue, relying on static map cues rather than dynamic grounding processes.

Read paper · arxiv.org → Multimodal Method Jun 29, 2026
Jun 29HF Daily Papers

GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity

Three seemingly distinct training methods for language models are shown to be variations of a single approach based on standard deviation adjustment, with the disagreement among sampled answers determining learning effectiveness and update…

Read paper · arxiv.org → Models Method Jun 29, 2026
Jun 29HF Daily Papers

HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

HealthAgentBench presents a comprehensive evaluation framework with 54 healthcare tasks across 7 categories to assess AI agents' capabilities in complex clinical workflows, revealing significant challenges in medical imaging and…

Read paper · arxiv.org → Science Benchmark Jun 29, 2026
Jun 29HF Daily Papers

Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming

AI-Infra-Guard is an open-source framework that addresses AI infrastructure security through layered detection paradigms spanning infrastructure, protocol, agent behavior, and model layers.

Read paper · arxiv.org → Safety System Jun 29, 2026
Jun 29HF Daily Papers

AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation

GraphRAG extends RAG by incorporating graph-structured data for LLMs, addressing latent feature misalignment through Adaptive-masking for Graph Embedding (AGE) that uses Transformer-based self-supervised learning with learnable node…

Read paper · arxiv.org → Safety Benchmark Jun 29, 2026
Jun 29HF Daily Papers

GORGO: Online Tuning for Cross-Region Network-Aware LLM Serving

GORGO is a proxy architecture that optimizes LLM inference load balancing by jointly considering network latency, prefill cost, and queueing delay through evolutionary strategy tuning on a new synthetic dataset.

Read paper · arxiv.org → Infra Dataset Jun 29, 2026
Jun 29HF Daily Papers

Teaching LLMs to Recommend and Defer in Underrepresented Epilepsy Care

A non-parametric prompt-learning framework called MANANA improves pediatric epilepsy treatment decisions by adapting to local prescribing practices and providing uncertainty-based deferral signals for low-confidence cases.

Read paper · arxiv.org → Evals System Jun 29, 2026
Jun 29HF Daily Papers

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

3D HAMSTER framework enhances robot manipulation by integrating a vision-language model with depth encoding to generate metrically accurate 3D trajectories for point cloud-based control policies.

Read paper · arxiv.org → Robotics System Jun 29, 2026
Jun 29HF Daily Papers

Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers

Cross-space distillation enables efficient knowledge transfer from high-capacity diffusion models to compact student models through a lightweight latent interface that aligns different VAE spaces.

Read paper · arxiv.org → Models Method Jun 29, 2026
Jun 29HF Daily Papers

Temporal Multi-Signal Fusion for Token-Level Hallucination Detection

Hallucination is detected as temporally extended spans via sequence labeling over fused external features, achieving robust cross-model performance without internal model access.

Read paper · arxiv.org → Evals Method Jun 29, 2026
Jun 28

CoffeeBench (Sakana AI)

Six LLM agents ran a simulated coffee-industry supply chain for 90 days: farmers, roasters, retailers, all negotiating prices and managing inventory. Large variance across frontier models. The notable failure was Claude Haiku 4.5, which went negative. Its reasoning logs showed coherent strategy throughout. It just kept choosing inaction. The gap between thinking and executing is the thing worth studying here.

Read paper · sakana.ai → Agents Analysis Jun 28, 2026
Jun 28

AllenAI on what hybrid models actually win at

Hybrid architectures (a few attention layers, recurrent layers for the rest) beat pure transformers on content words and pronouns. Pure transformers win on verbatim copying and bracket matching. The bigger point: a single aggregate loss number hides all of this. Comparing architectures on overall perplexity is the wrong instrument.

Read paper · huggingface.co → Infra System Jun 28, 2026
Jun 28

Beyond IID: How General Are Tabular Foundation Models, Really?

Tabular foundation models show varying performance across different data conditions, with traditional methods still outperforming newer approaches on complex, large-scale datasets.

Read paper · arxiv.org → Models Dataset Jun 28, 2026
Jun 28

Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark

The Nanotechnology Molecular Optimization (NMO) Benchmark introduces physics-based molecular design challenges that require new generative model approaches, moving beyond drug-discovery-focused metrics to enable scientific discovery in…

Read paper · arxiv.org → Science Benchmark Jun 28, 2026
Jun 28

The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction

ViDiHand uses pretrained video diffusion model representations with hand-overlay rendering to reconstruct 4D hand motion directly from video frames without detectors or optimization.

Read paper · arxiv.org → Multimodal Method Jun 28, 2026
Jun 28

DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model

DreamForge-World 0.1 Preview adapts a video generation architecture with a residual action pathway to enable real-time interactive world simulation on consumer hardware with low computational requirements.

Read paper · arxiv.org → Robotics Survey Jun 28, 2026
Jun 28

Walking in the Implicit: Interactive World Exploration via Neural Scene Representation

NeuWorld enables efficient interactive video generation by representing scenes as compact neural implicit states and using a transformer VAE with diffusion transformer for trajectory-conditioned rendering.

Read paper · arxiv.org → Multimodal Method Jun 28, 2026
Jun 28HF Daily Papers

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing

SafePyramid benchmark evaluates guardrail systems' ability to identify safety violations through in-context policy specification across multiple domains and complexity levels.

Read paper · arxiv.org → Safety Benchmark Jun 28, 2026
Jun 28HF Daily Papers

One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding

InnerZoom addresses GUI grounding challenges by preserving target-region awareness across decoder layers through a single-forward pass that bridges cross-layer evidence, achieving state-of-the-art performance with reduced computational…

Read paper · arxiv.org → Infra Method Jun 28, 2026
Jun 28HF Daily Papers

PoseShield: Neural Collision Fields for Human Self-Collision Resolution

PoseShield addresses self-collision issues in SMPL-based human pose estimation by applying neural collision constraints in pose space through constrained optimization and Eikonal regularization.

Read paper · arxiv.org → Models Method Jun 28, 2026
Jun 28HF Daily Papers

Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting

Flux-GS enables real-time high-fidelity 3D Gaussian Splatting on mobile platforms through efficient lighting representation, attribute-conditioned enhancement, and multi-view densification strategies.

Read paper · arxiv.org → Multimodal System Jun 28, 2026
Jun 28HF Daily Papers

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis

A masked discrete diffusion model for text-to-image synthesis that addresses limitations in token refinement and training efficiency through novel mechanisms and optimizations.

Read paper · arxiv.org → Multimodal Method Jun 28, 2026
Jun 28HF Daily Papers

TACO: Tool-Augmented Credit Optimization for Agentic Tool Use

Tool-Augmented Credit Optimization (TACO) improves multimodal agent performance by distinguishing useful, redundant, or misleading code operations through dual advantage channels: Differential Answer-Probe Reward for individual tool…

Read paper · arxiv.org → Multimodal Method Jun 28, 2026
Jun 28HF Daily Papers

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots

GUICrafter addresses GUI agent data challenges through a weakly-supervised approach using unannotated screenshots and a two-stage curriculum learning framework for visual grounding and reinforcement learning calibration.

Read paper · arxiv.org → Multimodal System Jun 28, 2026
Jun 28HF Daily Papers

Little Brains, Big Feats: Exploring Compact Language Models

Small language models can effectively perform retrieval-augmented generation tasks directly on-device without GPU acceleration.

Read paper · arxiv.org → Evals Benchmark Jun 28, 2026
Jun 28HF Daily Papers

PhotoQuilt: Training-Free Arbitrary-Resolution Photomosaics via Bootstrapped Tiled Denoising

PhotoQuilt is a training-free framework that generates high-resolution photomosaics by combining global layout composition with separate tile generation in latent space, overcoming limitations of diffusion models in balancing local detail…

Read paper · arxiv.org → Models System Jun 28, 2026
Jun 28HF Daily Papers

BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

BrainJanus represents the first unified brain model integrating brain, vision, and language through a shared Omni space, enabling bidirectional mapping between neural activity and sensory stimuli via a tokenized representation and…

Read paper · arxiv.org → Multimodal Method Jun 28, 2026
Jun 28HF Daily Papers

Orca: The World is in Your Mind

Orca establishes a unified world latent space through next-state-prediction modeling using multimodal data and demonstrates superior performance in downstream tasks compared to specialized baselines.

Read paper · arxiv.org → Multimodal Benchmark Jun 28, 2026
Jun 28HF Daily Papers

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

AVTok is a unified tokenizer for audio-video generation that uses a dual-stream transformer architecture with shared encoder-decoder and modal-specific queries to create compact one-dimensional latent representations.

Read paper · arxiv.org → Multimodal System Jun 28, 2026
Jun 28HF Daily Papers

DOPD: Dual On-policy Distillation

DOPD addresses privilege illusion in on-policy distillation by dynamically routing token-level supervision between teacher and student policies based on advantage gaps and probabilities, improving capability transfer in large and…

Read paper · arxiv.org → Multimodal Method Jun 28, 2026
Jun 28HF Daily Papers

LUMOS: A Semantic Operating-System Layer for Accessibility-Grounded AI Agents

LUMOS provides a semantic interaction layer that converts operating system metadata into machine-readable formats, enabling AI agents to interact more efficiently with computer interfaces than through traditional visual methods.

Read paper · arxiv.org → Multimodal System Jun 28, 2026
Jun 28HF Daily Papers

SWE-Together: Evaluating Coding Agents in Interactive User Sessions

SWE-Together is a multi-turn coding benchmark created from real user-agent interactions, featuring a reactive LLM simulator to evaluate agents based on both final correctness and interaction efficiency.

Read paper · arxiv.org → Evals Benchmark Jun 28, 2026
Jun 28HF Daily Papers

SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions

SWE-Interact presents a testbed that evaluates coding agents in realistic multi-turn, user-driven software engineering scenarios, revealing significant gaps between single-turn performance and interactive task completion.

Read paper · arxiv.org → Evals Benchmark Jun 28, 2026
Jun 28HF Daily Papers

Morphing into Hybrid Attention Models

FlashMorph is an efficient layer selection method that formulates hybrid layer selection as a budget-constrained optimization problem, using morphable models and linearization regularization to improve long-context efficiency in…

Read paper · arxiv.org → Infra Method Jun 28, 2026
Jun 28HF Daily Papers

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation

Scientific image generation faces challenges in semantic alignment and logical reasoning, prompting the creation of SciIR-82k dataset and SciIR-Bench evaluation framework to improve scientific reasoning capabilities in text-to-image models.

Read paper · arxiv.org → Multimodal Dataset Jun 28, 2026
Jun 28HF Daily Papers

CogSENet: Blind Image Deblurring with Blur-Conditioned Semantic Routing and Explicit Frequency Fusion

CogSENet presents a novel blind image deblurring framework inspired by eagle vision, incorporating semantic-aware modules and frequency decomposition for improved restoration quality and structural fidelity.

Read paper · arxiv.org → Multimodal System Jun 28, 2026
Jun 28HF Daily Papers

DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation

DuoMem is a dual-space distillation framework that transfers procedural problem-solving from large language models to compact student models through context-space and parameter-space distillation, achieving high performance with minimal…

Read paper · arxiv.org → Agents System Jun 28, 2026
Jun 28HF Daily Papers

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

A comprehensive benchmark called Musebench is introduced to evaluate multimodal large language models on nuanced artistic understanding, revealing a significant gap between current models and human expert performance in creative domain…

Read paper · arxiv.org → Multimodal Benchmark Jun 28, 2026
Jun 28HF Daily Papers

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications

A large-scale industrial platform leveraging LLMs and VLMs addresses key challenges in structured item knowledge production for e-commerce, achieving high precision and throughput across billions of products.

Read paper · arxiv.org → Infra System Jun 28, 2026
Jun 27

Prompt injection in AI résumé screening

Hiding self-promotional instructions inside résumé text moves candidates up in LLM rankings without adding real qualifications. The paper tests single and multi-injection settings. If you're using LLMs to filter applicants, the attack surface is open right now.

Read paper · arxiv.org → Models Method Jun 27, 2026
Jun 27

World model hallucinations aren't random

They concentrate in low-coverage regions of the state-action space, which makes them predictable. The paper shows they're also preventable. Directly relevant if you're building on generative world models for robotics or planning.

Read paper · arxiv.org → Robotics Method Jun 27, 2026
Jun 27HF Daily Papers

Learning Transferable Dynamics Priors from Action to World Modeling

Action-conditioned world modeling enables transferable dynamics priors for robot learning through pretraining on large-scale manipulation data, supporting both simulator-based policy evaluation and video-action prediction.

Read paper · arxiv.org → Robotics Benchmark Jun 27, 2026
Jun 27HF Daily Papers

Geometric Stability of Neural Population Codes: Regional Variation, Behavioral Relevance, and Circuit Dependence

Geometric stability measures the consistency of pairwise stimulus distances across trials, revealing a distinct aspect of neural representation that differs from temporal stability and decoding accuracy.

Read paper · arxiv.org → Evals Benchmark Jun 27, 2026
Jun 27HF Daily Papers

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

OSWorld 2.0 presents a comprehensive benchmark for evaluating computer-use agents through complex, real-world workflows that reveal current limitations in agent reasoning and task completion.

Read paper · arxiv.org → Evals Benchmark Jun 27, 2026
Jun 27HF Daily Papers

PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agents

POLICYGUARD is a sub-agent verifier that enhances LLM agent policy adherence by providing contextual reasoning and conversation-specific feedback across multi-turn interactions.

Read paper · arxiv.org → Agents Method Jun 27, 2026
Jun 27HF Daily Papers

Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views

A feed-forward framework decomposes 3D scenes into instance-structured token groups from multi-view images, enabling direct object-level reconstruction, segmentation, and manipulation without 3D annotations.

Read paper · arxiv.org → Robotics Dataset Jun 27, 2026
Jun 27HF Daily Papers

MirrorPPR: Exemplar-Based Portrait Photo Retouching

Exemplar-based portrait retouching framework using Diffusion Transformer with LoRA adaptation and self-augmented training data achieves superior quality and identity preservation.

Read paper · arxiv.org → Models System Jun 27, 2026
Jun 27HF Daily Papers

Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation

Vision-language dataset distillation method using rank-aware hyperbolic alignment to optimize synthetic image-text pairs for efficient contrastive model training while preserving modality-specific diversity.

Read paper · arxiv.org → Multimodal Dataset Jun 27, 2026
Jun 27HF Daily Papers

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

Training-inference mismatch in reinforcement learning for large language models leads to instability, which is addressed through a new policy optimization objective and framework that ensures consistent policy improvements between training…

Read paper · arxiv.org → Infra System Jun 27, 2026
Jun 27HF Daily Papers

SHAPE of Chain-of-Thought in Math Reasoning

SHAPE analyzes chain-of-thought reasoning via semantic spaces and heuristics to diagnose LLM mathematical reasoning and improve post-training.

Read paper · arxiv.org → Science Method Jun 27, 2026
Jun 26

DanceOPD

One diffusion model handling text-to-image, local editing, and global editing cleanly has been stubborn. Each capability tends to degrade the others during training, so most pipelines split them. This paper's on-policy generative field distillation approach trains one model to handle all three without the conflict. Builder read: watch for this pattern in the next wave of image-editing releases.

Read paper · arxiv.org → Multimodal System Jun 26, 2026
Jun 26

RiVER: RL without ground-truth answers

Standard RLVR needs correct answers to score model outputs, which limits it to tasks where ground truth exists. RiVER replaces that with ranking: compare outputs against each other and use relative quality as the reward signal. Builder read: useful if you're post-training on subjective tasks where "correct" is a judgment call.

Read paper · arxiv.org → Evals Method Jun 26, 2026
Jun 26HF Daily Papers

Agentic Abstention: Do Agents Know When to Stop Instead of Act?

Agentic abstention involves determining when an AI agent should cease interaction under uncertainty, requiring sequential decision-making across multiple environments and task types.

Read paper · arxiv.org → Agents Method Jun 26, 2026
Jun 26HF Daily Papers

Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks

Evolutionary fine-tuning enables large language models to develop cross-task problem-solving capabilities by learning from search trajectories, demonstrating improved performance on mathematical conjectures and optimization tasks.

Read paper · arxiv.org → Science Method Jun 26, 2026
Jun 26HF Daily Papers

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR

Transfer-Aware Curriculum (TAC) improves multi-domain reinforcement learning by prioritizing domains that provide broad benefits to other domains, using gradient-geometry alignment to estimate cross-domain transferability.

Read paper · arxiv.org → Safety Method Jun 26, 2026
Jun 26HF Daily Papers

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling

Sampling-based reasoning systems face a trade-off between coverage and selection, where additional samples beyond a few dozen provide diminishing returns and can degrade performance.

Read paper · arxiv.org → Infra System Jun 26, 2026
Jun 25

Real-time voice AI, tested on what it misses

GPT Realtime 2, Gemini 3.1 Flash Live, Qwen3.5 Omni Plus, Qwen Omni Flash all benchmarked on tasks where vocal delivery matters. All four miss paralinguistic cues (stress, hesitation, sarcasm). They hear the transcript, not the subtext. Real gap if you're shipping voice agents.

Evals Benchmark Jun 25, 2026
Jun 25

The unfireable safety kernel

System prompts, output filters, guardrail libraries all live inside the agent's own runtime — reachable by the agent. The paper's case: safety controls belong at the OS or infrastructure layer, outside the execution environment entirely. Right framing even if the implementations aren't there yet.

Robotics System Jun 25, 2026
Jun 25

GBC: Gradient-Based Connections for Optimizing Multi-Agent Systems

Gradient-Based Connections enables fine-grained attribution and optimization in multi-agent systems by modeling agent interactions as a computational graph and using gradient-based weights to identify error sources at the token level.

Read paper · arxiv.org → Infra System Jun 25, 2026
Jun 25

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots

Human manipulation skills are transferred to robots more effectively by using a bridging action representation based on relative wrist translation in the initial head-camera frame, combined with a vision-language-action model that handles…

Read paper · arxiv.org → Robotics Method Jun 25, 2026
Jun 25

ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering

A progressive multimodal search agent for knowledge-based visual question answering that adaptively selects search strategies and optimizes through sequence-level reinforcement learning.

Read paper · arxiv.org → Multimodal Method Jun 25, 2026
Jun 25

NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

Reinforcement learning post-training degrades perceptual quality in flow-based generators through velocity norm inflation, which requires training-time intervention rather than inference-time corrections to maintain both reward alignment…

Read paper · arxiv.org → Safety Method Jun 25, 2026
Jun 25

SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

SimFoundry enables zero-shot real-world robot policy training through automated simulation construction and diverse scene variations that improve generalization and performance prediction.

Read paper · arxiv.org → Robotics Benchmark Jun 25, 2026
Jun 25

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation

PhysisForcing enhances embodied video generation by enforcing physical consistency through pixel-level trajectory alignment and semantic-level relational alignment losses in a DiT-based framework.

Read paper · arxiv.org → Science System Jun 25, 2026
Jun 25HF Daily Papers

ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrieval

A fashion-specialized vision-language model achieves superior retrieval performance through full fine-tuning with knowledge distillation and weight interpolation, outperforming existing methods on a new benchmark while addressing…

Read paper · arxiv.org → Multimodal Benchmark Jun 25, 2026
Jun 25HF Daily Papers

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning

A new benchmark evaluates multimodal large language models' ability to reason over dynamic visual evidence through controlled temporal-logical operations rather than simple object recognition.

Read paper · arxiv.org → Robotics Benchmark Jun 25, 2026
Jun 25HF Daily Papers

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents

TUA-Bench presents a comprehensive benchmark for evaluating general-purpose terminal-use agents across diverse digital activities and specialized workflows, revealing significant performance gaps among current frontier agents.

Read paper · arxiv.org → Evals Benchmark Jun 25, 2026
Jun 25HF Daily Papers

ReFreeKV: Towards Threshold-Free KV Cache Compression

ReFreeKV addresses the limitations of threshold-dependent KV cache pruning by introducing a threshold-free approach that adaptively allocates compression budgets while maintaining full-cache performance across diverse datasets and model…

Read paper · arxiv.org → Models Dataset Jun 25, 2026
Jun 25HF Daily Papers

Dockerless: Environment-Free Program Verifier for Coding Agents

A Dockerless environment-free agentic patch verifier improves code patch evaluation accuracy and enables effective post-training without execution-based verification costs.

Read paper · arxiv.org → Evals Benchmark Jun 25, 2026
Jun 25HF Daily Papers

Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?

Research reveals that language backbones in Vision-Language-Action models are highly redundant for robotic manipulation tasks, while vision and action pathways are more critical, suggesting need for deliberate capacity allocation in future…

Read paper · arxiv.org → Robotics Method Jun 25, 2026
Jun 25HF Daily Papers

RocketSmith: Agentic Additive Manufacturing of High-Powered Rockets

An agentic system using large language models automates high-power rocket design processes, enabling successful flight testing with consistent simulation results.

Read paper · arxiv.org → Infra System Jun 25, 2026
Jun 25HF Daily Papers

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

PerceptionRubrics presents a rubric-based evaluation framework that identifies gaps between benchmark scores and real-world performance through atomic auditing and gated scoring mechanisms.

Read paper · arxiv.org → Multimodal Benchmark Jun 25, 2026
Jun 25HF Daily Papers

When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search

DiscoBench evaluates search agents' ability to handle ambiguous queries through clarification questioning and recovery in multi-step information-seeking tasks across diverse real-world domains.

Read paper · arxiv.org → Evals Benchmark Jun 25, 2026
Jun 25HF Daily Papers

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

Large Language Models fail to validate their outputs when evaluated through benchmarks, revealing a gap between task completion scores and actual implementation quality.

Read paper · arxiv.org → Evals Benchmark Jun 25, 2026
Jun 25HF Daily Papers

Parameter-Efficient Quantum-Inspired Fast Weight Programmers for Traffic-Matrix Forecasting

Quantum-inspired recurrent models using gated QKAN-FWPs demonstrate superior forecasting accuracy with reduced computational requirements compared to traditional LSTM networks for traffic matrix prediction.

Read paper · arxiv.org → Science Method Jun 25, 2026
Jun 25HF Daily Papers

DataComp-VLM: Improved Open Datasets for Vision-Language Models

DataComp for VLMs (DCVLM) establishes a comprehensive benchmark for evaluating data curation strategies in vision-language models, demonstrating that data mixing rather than filtering significantly improves model performance at scale.

Read paper · arxiv.org → Multimodal Dataset Jun 25, 2026
Jun 24

Grading the Grader evaluating an agentic data analysis system is not like scoring a quiz.

The key distinction the paper makes: a "diagnostic disagreement" (different explanation, same correct underlying finding) looks identical to a real error if your grader can't separate them. Anyone shipping automated evals for agents should read the taxonomy before they trust their numbers.

Read paper · arxiv.org → Evals Survey Jun 24, 2026
Jun 24

Paying to Know: agent-native micropayments in e-commerce

the core claim is that micropayment rails (x402, AP2) break the shopping-chatbot model. When the buyer is an autonomous agent that can independently verify specs, recommendations become cheap and abundant. What gets scarce is verified product data, and that's what commands a price.

Read paper · arxiv.org → Agents Method Jun 24, 2026
Jun 24

World Models in Pieces: structural certification for agents

formal argument that real-world agents are inevitably domain-specialized. The problem with current safety certification: a failure on a critical bottleneck looks identical to a failure on something irrelevant. The paper proposes structural certification that maps which failures actually matter. Useful framing for "is this agent safe to deploy" versus "did it pass the benchmark."

Read paper · arxiv.org → Safety Benchmark Jun 24, 2026
Jun 24

JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting

JetSpec is a speculative decoding framework that combines efficient forward drafting with causal conditioning to improve LLM inference speed and acceptance rates across various benchmarks.

Read paper · arxiv.org → Evals Benchmark Jun 24, 2026
Jun 24HF Daily Papers

In-Context World Modeling for Robotic Control

ICWM enables robot policies to infer system variables from self-generated interactions, allowing adaptation to novel configurations without parameter updates by treating system identification as an in-context adaptation problem.

Read paper · arxiv.org → Robotics System Jun 24, 2026
Jun 24HF Daily Papers

Confidence-Aware Tool Orchestration for Robust Video Understanding

Robust-TO addresses the Blind Trust Problem in video reasoning by integrating per-frame trustworthiness into an agentic framework that improves accuracy under realistic perturbations through calibrated evidence weighting and…

Read paper · arxiv.org → Multimodal System Jun 24, 2026
Jun 24HF Daily Papers

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

ViQ presents a visual quantization framework that balances semantic richness and detail preservation in discrete representations, enabling efficient multimodal training with native-resolution inputs.

Read paper · arxiv.org → Multimodal System Jun 24, 2026
Jun 24HF Daily Papers

OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

On-policy skill distillation framework extracts dense hindsight supervision from completed trajectories to improve language agent training efficiency and performance.

Read paper · arxiv.org → Multimodal System Jun 24, 2026
Jun 24HF Daily Papers

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

A unified agentic framework called Qwen-Image-Agent is proposed to address the context gap in text-to-image generation by progressively constructing complete generation context through planning, reasoning, searching, and memory mechanisms.

Read paper · arxiv.org → Multimodal System Jun 24, 2026
Jun 24HF Daily Papers

PhysiFormer: Learning to Simulate Mechanics in World Space

PhysiFormer uses coordinate-space diffusion to generate physically-plausible 3D object motions without explicit inductive biases, enabling efficient multi-object reasoning and generalization to complex materials and geometries.

Read paper · arxiv.org → Science Method Jun 24, 2026
Jun 24

Information-Aware KV Cache Compression for Long Reasoning

InfoKV is an entropy-aware KV cache compression framework that enhances long-context reasoning in LLMs by incorporating information-theoretic signals alongside attention weights.

Read paper · arxiv.org → Models System Jun 24, 2026
Jun 24

EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting

EO-WM is a video diffusion transformer for multispectral Earth Observation forecasting that incorporates physically informed conditioning frameworks to better capture weather-driven uncertainties in land-surface dynamics.

Read paper · arxiv.org → Multimodal System Jun 24, 2026
Jun 24

LISA: Likelihood Score Alignment for Visual-condition Controllable Generation

Score-based generative modeling reveals that side networks contribute likelihood scores to conditional control, leading to improved training efficiency through likelihood score alignment regularization.

Read paper · arxiv.org → Robotics Method Jun 24, 2026
Jun 24

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

A web-based benchmark evaluates agent generalization across challenging scenarios, revealing significant gaps between current agentic systems and human performance in temporal perception, graphical understanding, and 3D reasoning.

Read paper · arxiv.org → Multimodal Benchmark Jun 24, 2026
Jun 24

Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

A vision-language-action policy improved with reinforcement learning uses shared network predictions for success estimation and advantage calculation in bimanual garment folding, employing established RL techniques with novel optimization…

Read paper · arxiv.org → Multimodal Method Jun 24, 2026
Jun 24

Boundary-Aware Context Grounding for A Low-Channel EEG Agent

NeuraDock Agent combines a deterministic EEG processing engine with a language model interface to ensure accurate, hardware-aware analysis while maintaining local data security.

Read paper · arxiv.org → Safety System Jun 24, 2026
Jun 24HF Daily Papers

Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web Agents

A Korean web-agent benchmark evaluates breadth of search capabilities by requiring complete enumeration of entity memberships with attribute tables, revealing consistent failures in row recovery despite accurate set identification.

Read paper · arxiv.org → Evals Benchmark Jun 24, 2026
Jun 24HF Daily Papers

Qwen-Image-2.0-RL Technical Report

A reinforcement learning and on-policy distillation approach enhances the visual quality and instruction-following capabilities of a diffusion model for image generation and editing tasks.

Read paper · arxiv.org → Multimodal Analysis Jun 24, 2026
Jun 24HF Daily Papers

How Good Can Linear Models Be for Time-Series Forecasting?

Research demonstrates that preprocessing optimizations, particularly in context length, normalization, and regularization, can significantly improve time-series forecasting accuracy more effectively than scaling model architectures.

Read paper · arxiv.org → Infra System Jun 24, 2026
Jun 24HF Daily Papers

Large-Scale Tunnel Air-Ground Collaboration With FLISP: Fast LiDAR-IMU Synchronized Path Planner

Hydropower tunnel inspection is critical for infrastructure integrity yet remains inefficient and hazardous using manual methods. We propose FLISP (Fast LiDAR-IMU Synchronized Path Planner), a mapless planning framework for cooperative…

Read paper · arxiv.org → Agents System Jun 24, 2026
Jun 24HF Daily Papers

LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing

A novel streaming video editing framework enables causal, frame-by-frame editing with stable long-horizon preservation and real-time responsiveness through a three-stage distillation pipeline and AR-oriented mask cache.

Read paper · arxiv.org → Multimodal System Jun 24, 2026
Jun 24HF Daily Papers

Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE

SharpMoE addresses routing inefficiencies in diffusion models by using clean latent features to guide salient token identification and employs trajectory routing loss for precise compute allocation during multi-step denoising.

Read paper · arxiv.org → Models Method Jun 24, 2026
Jun 24HF Daily Papers

RedVox: Safety and Fairness Gaps in Speech Models Across Languages

Multilingual safety and fairness benchmark for speech models reveals persistent vulnerabilities across languages and naturalistic conditions.

Read paper · arxiv.org → Multimodal Benchmark Jun 24, 2026
Jun 24HF Daily Papers

PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh Generation

PolyFlow introduces a continuous mesh representation using a topology embedder and applies flow-matching with Transformers for parallel mesh generation, achieving faster inference and precise resolution control compared to autoregressive…

Read paper · arxiv.org → Robotics Method Jun 24, 2026
Jun 24HF Daily Papers

Delayed Verification Destabilizes Multi-Agent LLM Belief: Instability Thresholds and Optimal Corrector Placement

Delayed verification in multi-agent LLM systems can cause instability leading to oscillations, but grounded factual answering stabilizes the system by making truth an absorbing boundary.

Read paper · arxiv.org → Infra System Jun 24, 2026
Jun 23

SHERLOC

LLM coding agents spend roughly half their token budget just locating faults before writing a single edit. SHERLOC replaces naive file-retrieval localization with structured diagnostic output: it returns the fault site plus a diagnosis of why that location is the problem. That context is what a repair agent actually needs.

Read paper · arxiv.org → Evals Benchmark Jun 23, 2026
Jun 23

InSight

VLA robots are capped by their training demos: if a task wasn't in the dataset, the robot can't do it. InSight makes VLAs steerable at the primitive-action level, letting them chain novel behaviors from text instructions without new demonstrations. Autonomous skill acquisition from language alone, in a physical robot, is a meaningful step past imitation.

Read paper · arxiv.org → Robotics Dataset Jun 23, 2026
Jun 23

OpenThoughts-Agent

Most open agentic training efforts (SWE-Smith, SERA, Nemotron-Terminal) optimize for a single benchmark. This paper releases data recipes for training broadly capable agents across task types. The open release is the point. Leaderboard-specific training is why so many "agentic" models fall apart the moment you give them a real task.

Read paper · arxiv.org → Evals Benchmark Jun 23, 2026
Jun 23HF Daily Papers

Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints

Tool Suppression occurs when JSON Schema constraints and tool calling are jointly enabled, preventing open-weight models from invoking tools despite maintaining schema compliance, with the issue stemming from grammar-based token masking…

Read paper · arxiv.org → Agents Analysis Jun 23, 2026
Jun 23HF Daily Papers

MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation

A novel-view video synthesis method that enhances motion-aware diffusion models through multi-view point tracking supervision to improve geometric consistency and motion fidelity.

Read paper · arxiv.org → Multimodal Method Jun 23, 2026
Jun 23HF Daily Papers

ShutterMuse: Capture-Time Photography Guidance with MLLMs

Researchers developed a new benchmark and dataset for photography assistance, along with a unified multimodal model that provides both composition guidance and pose recommendations during image capture.

Read paper · arxiv.org → Multimodal Dataset Jun 23, 2026
Jun 23HF Daily Papers

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning

A novel label-free framework for visual reasoning called V-Zero is presented, which uses contrastive evidence gating to improve fine-grained visual reasoning without requiring annotated answer labels, achieving faster training than…

Read paper · arxiv.org → Multimodal System Jun 23, 2026
Jun 23HF Daily Papers

TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy

Camera-controllable video virtual try-on framework uses a 4D proxy with explicit human-environment decoupling and DiT-based video generation for omnidirectional viewing.

Read paper · arxiv.org → Robotics System Jun 23, 2026
Jun 23HF Daily Papers

Autodata: An agentic data scientist to create high quality synthetic data

Autodata enables AI agents to function as data scientists who create high-quality training data through meta-optimization, demonstrating improved performance across multiple task domains.

Read paper · arxiv.org → Agents Method Jun 23, 2026
Jun 23HF Daily Papers

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation

DomainShuttle enables open domain subject-driven text-to-video generation with high fidelity and flexibility across in-domain and cross-domain scenarios through domain-aware modeling and dual RoPE schemes.

Read paper · arxiv.org → Multimodal Method Jun 23, 2026
Jun 23HF Daily Papers

Improved Large Language Diffusion Models

Masked diffusion language models with fully bidirectional attention outperform autoregressive counterparts on various benchmarks while maintaining competitiveness with established models.

Read paper · arxiv.org → Evals Benchmark Jun 23, 2026
Jun 23HF Daily Papers

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models

Autoregressive video diffusion extends diffusion distillation frameworks to real-time streaming generation through causal training paradigms, achieving state-of-the-art performance with fast convergence and interactive world modeling…

Read paper · arxiv.org → Multimodal System Jun 23, 2026
Jun 23HF Daily Papers

The Verification Horizon: No Silver Bullet for Coding Agent Rewards

Verification challenges in AI agents arise from the difficulty of aligning proxy signals with human intent, requiring adaptive verification systems that evolve alongside generative capabilities.

Read paper · arxiv.org → Infra System Jun 23, 2026
Jun 23HF Daily Papers

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It

Research investigates how different supervisory signals and training strategies improve the stability and performance of large language models in tool-use tasks, addressing issues like catastrophic collapse and format sensitivity through…

Read paper · arxiv.org → Agents Analysis Jun 23, 2026
Jun 23HF Daily Papers

Fast LeWorldModel

Fast-LeWM accelerates visual planning by replacing autoregressive rollout with parallel action-prefix prediction, reducing computational costs and latency accumulation during long-horizon predictions.

Read paper · arxiv.org → Multimodal Method Jun 23, 2026
Jun 23HF Daily Papers

COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami

A computational origami system generates crease patterns from natural language using AI-driven optimization and aesthetic evaluation, enabling human-AI collaboration in mathematically constrained design.

Read paper · arxiv.org → Science Benchmark Jun 23, 2026
Jun 23HF Daily Papers

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation

A vision-language model-based hierarchical question graph framework evaluates video generation models' adherence to physical laws with granular violation detection and human correlation validation.

Read paper · arxiv.org → Science Benchmark Jun 23, 2026
Jun 23

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

Reinforcement learning post-training enables effective step-level scoring for language models without requiring dedicated reward model training by deriving an implicit advantage function called progress advantage.

Read paper · arxiv.org → Agents Method Jun 23, 2026
Jun 23HF Daily Papers

The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar

Tatoxa is a state-of-the-art text detoxification system for the Tatar language that demonstrates superior performance over existing LLMs and highlights the challenges of cross-lingual transfer in low-resource settings.

Read paper · arxiv.org → Infra System Jun 23, 2026
Jun 23HF Daily Papers

TheoremGraph: Bridging Formal and Informal Mathematics

A unified mathematical dependency graph connects informal and formal mathematics through semantic embedding and automated extraction from arXiv papers and Lean projects.

Read paper · arxiv.org → Science Method Jun 23, 2026
Jun 23HF Daily Papers

AI translation of literary texts is "fine", but readers still prefer human translations

Human readers prefer human-translated literary works over machine translations, finding the latter less immersive and harder to distinguish from human translations, despite machine translation metrics favoring the automated versions.

Read paper · arxiv.org → Evals Method Jun 23, 2026
Jun 23HF Daily Papers

Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?

A reinforcement learning framework called Play2Perfect enables sample-efficient robotic assembly tasks by first learning general manipulation skills through playful interaction with diverse objects, then adapting these skills for precise…

Read paper · arxiv.org → Robotics System Jun 23, 2026
Jun 22Editor's picks

Sovereign Execution Brokers authorizing an agent's identity isn't the same as authorizing its actions.

This paper proposes a broker layer that sits between agentic reasoning and production mutation calls, requiring certificate-bound authority per action. An agent with a valid IAM token can still hallucinate a destructive DELETE. Identity is necessary, not sufficient.

Read paper · arxiv.org → Agents Method Jun 22, 2026
Jun 22Editor's picks

Hierarchical Recovery for Cross-Device Agents

when an agent fails mid-task (browser times out, API drops), the standard response is full replanning. Expensive and usually unnecessary. Local recovery at the failure point outperforms global replanning on both success rate and latency. If you're building multi-device orchestration, subtask-level fault handling beats a full restart.

Read paper · arxiv.org → Agents Method Jun 22, 2026
Jun 22Editor's picks

How Transparent is DiffusionGemma?

diffusion LMs generate in a continuous latent space, meaning more computation happens somewhere you can't read. This paper finds that does hurt interpretability relative to autoregressive models. If "why did it do that" matters for your use case, that's a real cost to weigh against the speed argument.

Read paper · arxiv.org → Models Method Jun 22, 2026
Jun 22HF Daily Papers

What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics

Jailbreak attacks expose vulnerabilities in aligned large language models, revealing that harmful intent is encoded in structured intermediate uncertainty dynamics rather than output representations.

Read paper · arxiv.org → Safety Method Jun 22, 2026
Jun 22HF Daily Papers

CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression

Two-channel evaluation shows output compression reduces costs while input compression increases costs and degrades accuracy across models and datasets.

Read paper · arxiv.org → Evals Dataset Jun 22, 2026
Jun 22HF Daily Papers

Are We Ready For An Agent-Native Memory System?

Large language model agents' memory systems have evolved into complex data management frameworks requiring systematic evaluation across multiple modules and workloads to understand their performance characteristics and trade-offs.

Read paper · arxiv.org → Evals Survey Jun 22, 2026
Jun 22HF Daily Papers

RoPE-Aware Bit Allocation for KV-Cache Quantization

Block-GTQ introduces a RoPE-aware bit allocation method for key-cache quantization that improves attention accuracy and downstream performance through adaptive bit distribution and packed cache serving.

Read paper · arxiv.org → Infra Method Jun 22, 2026
Jun 22HF Daily Papers

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation

Implicit Visual Chain-of-Thought decomposes visual conditioning into structural and semantic cascades for improved structure-aware image generation with sketch supervision.

Read paper · arxiv.org → Multimodal Method Jun 22, 2026
Jun 22HF Daily Papers

Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods

A large-scale synthetic dataset and specialized model architecture are introduced to address the challenges of artistic text recognition by improving data diversity and model flexibility for irregular text layouts.

Read paper · arxiv.org → Infra Dataset Jun 22, 2026
Jun 22HF Daily Papers

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

Wan-Streamer is a unified, end-to-end multimodal model that enables real-time audio-visual interaction through causal attention mechanisms and integrated processing of visual, audio, and text modalities.

Read paper · arxiv.org → Multimodal Method Jun 22, 2026
Jun 22HF Daily Papers

MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery

Long-term memory in LLM agents should be evaluated as an auditable post-interaction artifact by reconstructing structured user state from the agent's memory, as demonstrated by MEMPROBE, a benchmark testing memory recovery against…

Read paper · arxiv.org → Evals Benchmark Jun 22, 2026
Jun 22HF Daily Papers

Do Thinking Tokens Help with Safety?

Research reveals that reasoning models' safety outcomes are predictable from early hidden representations, with deliberation appearing but not substantially influencing final responses, and current safety interventions inadvertently…

Read paper · arxiv.org → Safety Method Jun 22, 2026
Jun 22HF Daily Papers

Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matching

Lite Any Stereo V2 (LAS2) presents an efficient stereo matching approach that achieves state-of-the-art accuracy with significantly reduced latency through optimized architecture and training strategies.

Read paper · arxiv.org → Infra System Jun 22, 2026
Jun 22HF Daily Papers

Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents

Conversational infill enables small real-time models to maintain responsiveness while integrating delayed reasoning outputs, bridging the gap between latency and capability in voice agents.

Read paper · arxiv.org → Agents Method Jun 22, 2026
Jun 22HF Daily Papers

Trimming the Long-Tail of Visual World Modeling Evaluation

Current visual world models demonstrate limited generalization beyond common physical interactions, struggling with rare and irregular scenarios despite achieving realism on standard benchmarks.

Read paper · arxiv.org → Multimodal Benchmark Jun 22, 2026
Jun 22HF Daily Papers

SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History

SkillHone enables continuous evolution of agent skills by maintaining persistent decision histories and incorporating practice feedback for improved performance across research and tool-mediated analysis tasks.

Read paper · arxiv.org → Agents Analysis Jun 22, 2026
Jun 22HF Daily Papers

LLM Program Optimization via Retrieval Augmented Search

Blackbox adaptation methods using retrieval-augmented search and atomic edit decomposition improve program optimization performance for both C++ and Python code.

Read paper · arxiv.org → Evals Benchmark Jun 22, 2026
Jun 21Editor's picks

Probe-and-Refine Tuning for Coding Agents

an automated method for improving AGENTS.md files from the bottom up. The agent probes the codebase, identifies where its own understanding breaks down, then rewrites the guidance doc. Tested on SWE-bench-style evals; improved files produced measurably better task success. If your coding agent keeps getting the same subsystem wrong, the file is the problem.

Read paper · arxiv.org → Evals Benchmark Jun 21, 2026
Jun 21Editor's picks

LLM Psychological Profiles Are Largely Measurement Artifacts

formal psychometric analysis applied to the tools used to assign LLMs stable "personality" scores found the profiles weren't consistent in ways that would make them meaningful. Product teams using personality benchmarks to evaluate or fine-tune model character should treat those numbers skeptically until somebody builds instruments designed for LLMs rather than for humans.

Read paper · arxiv.org → Evals Benchmark Jun 21, 2026
Jun 21HF Daily Papers

ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation

ReNIO enhances on-policy distillation for language models by reweighting negative trajectories based on token-level probability ratios, improving reasoning performance in mathematical and code generation tasks.

Read paper · arxiv.org → Science Method Jun 21, 2026
Jun 21HF Daily Papers

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

The book provides a comprehensive guide to building autonomous AI systems, covering foundational elements like transformer architecture and training methods, along with advanced topics such as reinforcement learning, agent architectures,…

Read paper · arxiv.org → Infra System Jun 21, 2026
Jun 21HF Daily Papers

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct

A novel framework called VeriEvol is introduced that addresses the challenge of scaling reinforcement learning for visual mathematical reasoning by ensuring reliable reward labels through a two-axis approach that separates prompt…

Read paper · arxiv.org → Science System Jun 21, 2026
Jun 21HF Daily Papers

Critique of Agent Model

True artificial agency requires internalized structures for goals, identity, decision-making, self-regulation, and learning, distinguishing autonomous systems from task-specific ones.

Read paper · arxiv.org → Infra System Jun 21, 2026
Jun 21HF Daily Papers

GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents

Computer-use agents can execute software tasks through either graphical interfaces or programmatic command interfaces, but existing evaluations confound interaction modality with differences in tasks, initial states, verifiers, and…

Read paper · arxiv.org → Evals Benchmark Jun 21, 2026
Jun 21HF Daily Papers

Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents

Standard LLM agents rely on plan content remaining in context rather than maintaining it as persistent state, with evidence shown through replay pairing diagnostics and compression stress tests.

Read paper · arxiv.org → Agents Method Jun 21, 2026
Jun 21

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

ABACUS is a unified vision-language model that performs object counting and related tasks through innovative spatial grounding, boundary-aware counting policies, and self-critical learning strategies.

Read paper · arxiv.org → Multimodal Method Jun 21, 2026
Jun 21HF Daily Papers

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning

SingGuard is a policy-adaptive multimodal guardrail system that evaluates safety in real-time conversations by dynamically applying natural-language rules through fast-to-slow reasoning modes.

Read paper · arxiv.org → Multimodal Benchmark Jun 21, 2026
Jun 21HF Daily Papers

ReasoningLens: Hierarchical Visualization and Diagnostic Auditing for Large Reasoning Models

ReasoningLens is an open-source framework that provides hierarchical visualization and diagnostic auditing for complex reasoning chains in large reasoning models, enabling structured analysis and error detection through interactive…

Read paper · arxiv.org → Multimodal System Jun 21, 2026
Jun 21HF Daily Papers

Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation

Procedural memory enhances LLM agents on workplace tasks through skill transfer across roles and models, with varying generalization capabilities affecting deployment strategies.

Read paper · arxiv.org → Robotics Benchmark Jun 21, 2026
Jun 21HF Daily Papers

Mind the Heads: Topological Representation Alignment for Multimodal LLMs

HeRA aligns individual attention heads in MLLMs to preserve local neighborhood relationships across modalities, improving vision-centric task performance and reducing visual hallucinations.

Read paper · arxiv.org → Multimodal Method Jun 21, 2026
Jun 21HF Daily Papers

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos

Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand systems for lack of large-scale, language-aligned, and action-accurate demonstration data. To address this bottleneck, we…

Read paper · arxiv.org → Robotics System Jun 21, 2026
Jun 21HF Daily Papers

TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs

Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and training the model to generate it via constrained beam search. Toolsense shows that this regime has two critical drawbacks: it…

Read paper · arxiv.org → Evals Benchmark Jun 21, 2026
Jun 20Editor's picks

Contagion Networks bias doesn't stay local in multi-agent systems.

A new framework formalizes how evaluator biases spread when LLMs serve as judges inside multi-agent pipelines. A controlled three-agent experiment confirmed the propagation is fast and measurable. If you're using an LLM as a judge anywhere in your pipeline, audit it like the root of a tree, not a leaf.

Read paper · arxiv.org → Robotics Benchmark Jun 20, 2026
Jun 20Editor's picks

LedgerAgent instead of restuffing task state into the prompt on every turn, this paper proposes typed ledger objects that carry facts, constraints, and domain conditions across a conversation.

Cleaner than context stuffing, and auditable by design. Worth reading for anyone building multi-turn tool-calling flows with policy requirements.

Read paper · arxiv.org → Infra Method Jun 20, 2026
Jun 20HF Daily Papers

Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do

Multimodal Chain-of-Thought reasoning shows selective effectiveness across different tasks, with limitations in maintaining visual introspection during reasoning processes.

Read paper · arxiv.org → Multimodal Method Jun 20, 2026
Jun 20HF Daily Papers

Interleaved Speech Language Models Latently Work In Text

Interleaved speech-text language models exhibit an implicit transcription phase where text tokens become decodable in intermediate layers, followed by text-based prediction before speech domain transformation.

Read paper · arxiv.org → Multimodal Method Jun 20, 2026
Jun 19Editor's picks

MosaicLeaks (ServiceNow) tests whether research agents leak proprietary information while completing assigned tasks.

They do, more than expected. Required reading before deploying agents against internal knowledge bases.

Read paper · huggingface.co → Agents Method Jun 19, 2026
Jun 19Editor's picks

Beyond LoRA is Hugging Face's systematic comparison of PEFT fine-tuning methods, finding several alternatives that outperform LoRA in specific settings.

If you haven't revisited fine-tuning choices recently, this is the reason to.

Read paper · huggingface.co → Infra System Jun 19, 2026
Jun 19HF Daily Papers

EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies

EBench is a comprehensive simulation benchmark for evaluating generalist mobile manipulation policies across diverse tasks and dimensions, revealing distinct capability profiles and generalization patterns among state-of-the-art models.

Read paper · arxiv.org → Robotics Benchmark Jun 19, 2026
Jun 19HF Daily Papers

Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation

Multi4D addresses the trade-off between motion consistency and visual fidelity in dynamic 3D Gaussian splatting through a multi-level competitive allocation framework that enables adaptive specialization and efficient representation.

Read paper · arxiv.org → Multimodal System Jun 19, 2026
Jun 19HF Daily Papers

OpenBioRQ: Unsolved Biomedical Research Questions for Agents

A new biomedical benchmark evaluates agentic models' ability to verify sources and avoid false citations by testing unsolved research questions with no answer keys, revealing significant failures in retrieval-grounded reasoning and tool…

Read paper · arxiv.org → Science Benchmark Jun 19, 2026
Jun 18Editor's picks

Google AMIE in Nature

Google's medical AI system matched primary care physicians in managing complex, multi-condition diseases over extended patient consultations, per a peer-reviewed study published in Nature. The comparison covered real-world chronic disease management scenarios across a range of conditions. This is not a toy benchmark.

Read paper · blog.google → Science Survey Jun 18, 2026
Jun 18Editor's picks

Near-autonomous AI chemist

OpenAI and Molecule.one used GPT-5.4 to systematically improve Chan-Lam coupling, a drug-synthesis reaction where yields have historically been too low for practical pharmaceutical use. The AI proposed experimental directions; Molecule.one ran synthesis experiments; human chemists validated representative results and confirmed yield improvements across a majority of compounds tested. Full cycle: about 2.5 months.

Read paper · openai.com → Infra System Jun 18, 2026
Jun 18Editor's picks

Data Intelligence Agents

(arXiv). A system of three agents (Data Interpreter, Schema Creator, Query Agent) that automates the handoff between data owners, engineers, and analysts in enterprise pipelines. The central claim: the biggest bottleneck in production data integration is repeated, lossy translation between roles, and coding agents can close most of it.

Read paper · arxiv.org → Infra System Jun 18, 2026
Jun 18HF Daily Papers

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

UnityShots is a memory-driven audio-video generation system that maintains consistent subject appearance and audio across video cuts using fixed-size long-term and short-term memory slots with boundary-conditioned gates and discrete…

Read paper · arxiv.org → Multimodal System Jun 18, 2026
Jun 18HF Daily Papers

PrivacyAlign: Contextual Privacy Alignment for LLM Agents

Researchers develop a human-centered approach to align AI agents with privacy norms by creating a comprehensive dataset of privacy judgments and using annotation-conditioned reward modeling to improve agent behavior.

Read paper · arxiv.org → Safety Dataset Jun 18, 2026
Jun 18HF Daily Papers

Discretizing Reward Models

Reward models in reinforcement learning suffer from oversensitivity issues where they assign different scores to equally good responses, leading to poor policy learning, but this can be mitigated through discretization techniques that…

Read paper · arxiv.org → Models Method Jun 18, 2026
Jun 18HF Daily Papers

Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

A novel speaker verification framework combines frozen self-supervised features with ECAPA-TDNN and MoE modules to improve identity verification across both speech and non-verbal vocalizations while maintaining speech performance.

Read paper · arxiv.org → Multimodal System Jun 18, 2026
Jun 18HF Daily Papers

BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discovery

BioInsight is a multi-agent system that transforms static biomedical reports into interactive, evidence-centered interfaces by organizing disease-specific evidence through structured artifacts and deterministic citation normalization.

Read paper · arxiv.org → Science System Jun 18, 2026
Jun 17Editor's picks

Red-teaming Anthropic Fable 5 and Opus 4.8

independent researchers tested both models against 7,826 harmful intents across 10 harm categories using the HackAgent automated jailbreak framework. The paper maps which attack families succeed and at what rates. Useful baseline if you're building on either model and need to understand the residual attack surface.

Read paper · arxiv.org → Safety Benchmark Jun 17, 2026
Jun 17Editor's picks

All Smoke, No Alarm

an audit of 932,000+ AI-agent-authored pull requests across 116,000 GitHub repos found that test files routinely lack oracle signals. Assertions, expected values, anything that would actually catch a failure. Agents write the test skeleton; the verification logic often isn't there. Worth checking your own codegen pipeline.

Read paper · arxiv.org → Agents System Jun 17, 2026
Jun 17Editor's picks

Structural Role Injection in Handlebars-Templated LLM Prompts

Handlebars' triple-brace {{{x}}} syntax skips HTML escaping and is the default in Microsoft Semantic Kernel. The paper shows this creates a structural injection vector where malicious input can redefine roles within the prompt. If you're using templated prompts, default to double-brace and audit any {{{ usage in your codebase.

Read paper · arxiv.org → Models Method Jun 17, 2026
Jun 17HF Daily Papers

Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation

DO-ALL is a test-time adaptation framework that uses dataset distillation to create synthetic anchors for stable long-term model performance without retaining source data.

Read paper · arxiv.org → Models Dataset Jun 17, 2026
Jun 17HF Daily Papers

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

LLM agents frequently select higher-privilege tools unnecessarily, and while safety alignment doesn't ensure least-privilege choices, a post-training defense can reduce excessive privilege use without sacrificing performance.

Read paper · arxiv.org → Safety Method Jun 17, 2026
Jun 17HF Daily Papers

Hadith computational science in the age of large language models: a critical narrative review

Hadith computational science is evaluated as an evidence infrastructure challenge requiring integration of transformer and retrieval-based methods with expert validation and provenance.

Read paper · arxiv.org → Science Survey Jun 17, 2026
Jun 16Editor's picks

TokenPilot: cache-efficient context management for LLM agents

Most KV cache pruning approaches break prefix matching by mutating token sequences; TokenPilot keeps the prefix stable while still compressing context, cutting recompute cost on long sessions. The right kind of infrastructure paper for anyone running multi-step agents.

Read paper · arxiv.org → Agents Method Jun 16, 2026
Jun 16Editor's picks

ContextRL: teaching agents to notice the decisive detail

LLMs reliably miss the one critical line in a long tool trace or the subtle clue in an image. ContextRL adds a context-awareness objective to the RL loop so the model learns to surface that detail rather than gloss over it. Direct improvement for any agent reading external tool output.

Read paper · arxiv.org → Multimodal Method Jun 16, 2026
Jun 16HF Daily Papers

Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement

An object-centric residual reinforcement learning framework improves real-world vision-language-action model robustness through simulation-trained corrective policies that transfer zero-shot despite sim-to-real challenges.

Read paper · arxiv.org → Multimodal System Jun 16, 2026
Jun 16HF Daily Papers

TurboServe: Serving Streaming Video Generation Efficiently and Economically

TurboServe is a specialized serving system for streaming video generation that addresses session state management and dynamic resource allocation challenges through integrated scheduling, autoscaling, and migration mechanisms.

Read paper · arxiv.org → Multimodal System Jun 16, 2026
Jun 15Editor's picks

When verifiers backfire

A paper out today shows that verifier-driven self-DPO, the standard recipe for self-improving production VLMs, can cause regression on tasks outside the verifier's training distribution. Scores go up on seen tasks while performance quietly drops on unseen ones. If you're running RLVR-based self-improvement in production, this is the failure mode to test before your next rollout.

Read paper · arxiv.org → Models Method Jun 15, 2026
Jun 15Editor's picks

Reasoning for live inputs

AdaSR proposes a model that reasons as context arrives rather than after the full input is in. The target: live audio, video streams, continuous sensor feeds. Still a research paper, but the problem it names ("read-then-think doesn't fit dynamic input") is exactly the gap real-time agent work keeps running into.

Read paper · arxiv.org → Multimodal Method Jun 15, 2026
Jun 15HF Daily Papers

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence

This survey explores multimodal code intelligence systems that generate and reason with code based on visual inputs, categorizing approaches across GUI, scientific visualization, structured graphics, and emerging frameworks while…

Read paper · arxiv.org → Multimodal Survey Jun 15, 2026
Jun 14Editor's picks

AgentSpec

Controlled composition experiments on embodied agent scaffolds, isolating what each component (memory, reflection, action execution, learning) actually contributes when tested independently. Adding more components is not additive: some combinations improve performance; some degrade it. Start lean, add one piece with a real eval before the next.

Read paper · arxiv.org → Robotics Benchmark Jun 14, 2026
Jun 14Editor's picks

SIMMER

A benchmark for latent failures in LLM-based sequential planning. These are plan steps that look valid but only fail several actions later, after the world state has changed. Existing benchmarks miss them entirely. Worth reading if you're designing recovery or rollback logic for multi-step agents.

Read paper · arxiv.org → Evals Benchmark Jun 14, 2026
Jun 14HF Daily Papers

RL-Index: Reinforcement Learning for Retrieval Index Reasoning

RL-Index introduces an agentic indexing framework that shifts reasoning from query time to indexing stage by using LLM-generated rationales and reinforcement learning to improve retrieval effectiveness and reduce latency.

Read paper · arxiv.org → Evals Benchmark Jun 14, 2026
Jun 14HF Daily Papers

CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies

CoffeeBench evaluates LLM agents in a multi-agent economic simulation where firms interact over 90 days to maximize profits, revealing differences in communication patterns and performance among various models.

Read paper · arxiv.org → Evals Benchmark Jun 14, 2026
Jun 14HF Daily Papers

How Post-Training Shapes Biological Reasoning Models

Post-training stages in biological reasoning models differently affect generalization, with continued pre-training aligning models with biological language, supervised fine-tuning improving in-domain performance but reducing out-of-domain…

Read paper · arxiv.org → Models Method Jun 14, 2026
Jun 11Editor's picks

Doc-to-Atom: Compiling and Composing Memory Atoms

Proposes turning long documents into reusable "memory atoms" so models sidestep attention's quadratic cost in multi-turn reasoning, cutting the memory and latency tax on long contexts.

Read paper · arxiv.org → Agents Method Jun 11, 2026
Jun 11Editor's picks

TAHOE: Text-to-SQL with Automated Hint Optimization

Tackles why Text-to-SQL prototypes fail in production by auto-tuning query hints from past runs to handle strict SQL dialects at scale. Useful reading if you're putting natural-language-to-DB anywhere near real data.

Read paper · arxiv.org → Models Method Jun 11, 2026
Jun 10HF Daily Papers

On Subquadratic Architectures: From Applications to Principles

xLSTM demonstrates superior performance in sequence modeling tasks compared to Mamba-2 and Gated DeltaNet due to enhanced state tracking and memory dynamics.

Read paper · arxiv.org → Infra System Jun 10, 2026
Jun 10HF Daily Papers

Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training

ART enables parameter-efficient fine-tuning of frozen multimodal language models by optimizing raw visual input through gradient backpropagation, achieving performance comparable…

Read paper · arxiv.org → Multimodal Method Jun 10, 2026
Jun 10HF Daily Papers

HarmProfile: Characterizing Harmful Distributions in Frontier LLMs

HarmProfile is a benchmark dataset that characterizes frontier LLM safety failures through content analysis, revealing that harmfulness and diversity increase with model capability.

Read paper · arxiv.org → Safety Dataset Jun 10, 2026
Jun 9HF Daily Papers

Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs

ModSleuth is an agentic system that recursively reconstructs large-scale dependency graphs for LLM development by analyzing public artifacts and resolving inconsistencies in docum…

Read paper · arxiv.org → Infra System Jun 9, 2026
Jun 9HF Daily Papers

APEX: A Network-Native Time-Series Foundation Model for Forecasting and Anomaly Detection for Wireless Edge Operations

Network-native transformer model APEX demonstrates superior forecasting performance for wireless network telemetry compared to existing foundation models and traditional methods.

Read paper · arxiv.org → Evals Method Jun 9, 2026
Jun 9HF Daily Papers

Towards Diverse Scientific Hypothesis Search with Large Language Models

Evolutionary framework for hypothesis generation that improves diversity and quality through multi-temperature sampling and information exchange across search levels.

Read paper · arxiv.org → Models System Jun 9, 2026
Jun 9HF Daily Papers

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models

Vision-language models can improve grounding performance under aggressive token reduction by replacing irreversible visual-token pruning with recoverable routing that allows token…

Read paper · arxiv.org → Multimodal Method Jun 9, 2026
Jun 9HF Daily Papers

Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models

SKIM is an adaptive multi-resolution soft token compression framework that efficiently compresses procedural skills while maintaining task performance and enabling lightweight off…

Read paper · arxiv.org → Models System Jun 9, 2026
Jun 9HF Daily Papers

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

TRACE is a rollout allocation framework that improves reward contrast in multi-turn agentic reinforcement learning by dynamically distributing resources across tree-structured rol…

Read paper · arxiv.org → Agents System Jun 9, 2026
Jun 9HF Daily Papers

Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code

Grammar-constrained decoding techniques used to ensure syntactic validity in code generation can be exploited as an attack surface, leading to the development of a jailbreak metho…

Read paper · arxiv.org → Safety Method Jun 9, 2026
Jun 9HF Daily Papers

Time-Series Foundation Model Embeddings for Remaining Useful Life Estimation

A lightweight approach combining a frozen pretrained time-series foundation model with a simple regression head achieves superior RUL prediction performance compared to various ba…

Read paper · arxiv.org → Models Method Jun 9, 2026
Jun 9HF Daily Papers

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning

A training-free framework for spatial reasoning from egocentric videos that enables revisiting conclusions through synthesized novel-view videos generated from predicted 3D geomet…

Read paper · arxiv.org → Multimodal System Jun 9, 2026
Jun 9HF Daily Papers

Redesign Mixture-of-Experts Routers with Manifold Power Iteration

Researchers propose a novel router redesign for Mixture-of-Experts models that aligns router rows with the principal singular directions of expert matrices using Manifold Power It…

Read paper · arxiv.org → Models Method Jun 9, 2026
Jun 9HF Daily Papers

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks

A new benchmark and adapter protocol called Claw-SWE-Bench enables fair comparison of diverse coding agents by standardizing evaluation conditions and revealing the importance of…

Read paper · arxiv.org → Evals Benchmark Jun 9, 2026
Jun 9HF Daily Papers

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement

An AI framework called Arbor enables autonomous scientific research by combining strategic coordination, isolated hypothesis testing, and a persistent knowledge tree to iterativel…

Read paper · arxiv.org → Models System Jun 9, 2026
Jun 9HF Daily Papers

Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization

Recursive automated composition framework enables scalable reinforcement learning for language models by automatically combining verifiable environments through compositional oper…

Read paper · arxiv.org → Models System Jun 9, 2026
Jun 9HF Daily Papers

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application

Large language model agents require specialized environments for training and evaluation, which can be categorized by their engineering lifecycle stages and evolved through variou…

Read paper · arxiv.org → Evals Survey Jun 9, 2026
Jun 8HF Daily Papers

ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatorics

A new benchmark called ComBench is introduced to evaluate large language models' combinatorial reasoning abilities through Olympiad-level problems that test both proof constructio…

Read paper · arxiv.org → Evals Benchmark Jun 8, 2026
Jun 8HF Daily Papers

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

Embodied-R1.5 is a unified embodied foundation model that integrates embodied reasoning capabilities and achieves state-of-the-art performance on embodied vision-language benchmar…

Read paper · arxiv.org → Robotics Method Jun 8, 2026
Jun 8HF Daily Papers

Forecasting Future Behavior as a Learning Task

Behavior Forecasters are trained to predict large reasoning model outputs from single trajectories, outperforming large language models while requiring significantly less computational cost.

Read paper · arxiv.org → Models Method Jun 8, 2026
Jun 7HF Daily Papers

τ-Rec: A Verifiable Benchmark for Agentic Recommender Systems

A benchmark for agentic recommender systems is introduced that uses verifiable rewards and controlled dialogue constraints to evaluate conversational agent reliability, revealing…

Read paper · arxiv.org → Robotics Benchmark Jun 7, 2026
Jun 7HF Daily Papers

FlowLet: Conditional 3D Brain MRI Synthesis using Wavelet Flow Matching

FlowLet is a conditional generative framework that synthesizes age-conditioned 3D MRIs using flow matching in an invertible 3D wavelet domain, improving brain age prediction perfo…

Read paper · arxiv.org → Multimodal System Jun 7, 2026
Jun 7HF Daily Papers

TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders

TRL-Bench establishes a standardized benchmark for evaluating tabular representation learning models across multiple granularities, revealing that encoder performance varies by ta…

Read paper · arxiv.org → Evals Benchmark Jun 7, 2026
Jun 7HF Daily Papers

Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions

A teacher-student framework decouples complex reasoning from efficient reward deployment in text-to-image training, achieving superior preference accuracy and optimization perform…

Read paper · arxiv.org → Multimodal System Jun 7, 2026
Jun 6Editor's picks

NVIDIA at CVPR: physical-AI "agent skills."

New building blocks that wrap a full workflow around robotics, autonomous-vehicle, and vision models. The bet: the bottleneck in physical AI isn't stronger models but the tooling around them.

Read paper · blogs.nvidia.com → Robotics Method Jun 6, 2026
Jun 5HF Daily Papers

POISE: Position-Aware Undetectable Skill Injection on LLM Agents

POISE is a stealthy skill-poisoning attack that embeds malicious triggers within benign-looking instructions, achieving high attack success rates while avoiding detection by LLM s…

Read paper · arxiv.org → Evals Method Jun 5, 2026
Jun 4HF Daily Papers

ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction

ReVision improves computer-use agent efficiency by removing redundant visual patches from consecutive screenshots while preserving spatial structure, reducing token usage by 46% a…

Read paper · arxiv.org → Multimodal Method Jun 4, 2026
Jun 4HF Daily Papers

Breaking the Bubble: Asynchronous Pipeline Parallel Training with Bounded Weight Inconsistency

PACI enables efficient asynchronous pipeline training by controlling forward/backward weight inconsistency through local gradient accumulation, achieving higher throughput and fas…

Read paper · arxiv.org → Robotics System Jun 4, 2026
Jun 3HF Daily Papers

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models

DRIFT is a framework that adapts pretrained vision-language models for continuous decoding tasks by combining coarse prediction with iterative refinement through flow matching, im…

Read paper · arxiv.org → Multimodal System Jun 3, 2026
Jun 2HF Daily Papers

SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference

SparDA is a decoupled sparse attention architecture that improves long-context LLM inference by reducing KV cache bottlenecks and attention complexity through aForecast projection…

Read paper · arxiv.org → Infra System Jun 2, 2026
Jun 1HF Daily Papers

Large Language Models Are Overconfident in Their Own Responses

Instruction tuning degrades calibration in large language models, with chat templates exacerbating overconfidence through ownership bias, which can be mitigated by reframing model…

Read paper · arxiv.org → Models Method Jun 1, 2026
Jun 1HF Daily Papers

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning

EvoTrainer autonomously evolves both language model policies and training harnesses through empirical feedback, demonstrating superior performance in complex reasoning and coding…

Read paper · arxiv.org → Agents Method Jun 1, 2026
Jun 1HF Daily Papers

GridVQA-X: A Framework for Evaluating Multimodal Explainability Methods

GridVQA-X introduces a diagnostic framework to evaluate cross-modal explainability by distinguishing genuine spatial-relational reasoning from cross-modal shortcuts in multimodal models.

Read paper · arxiv.org → Multimodal Benchmark Jun 1, 2026