Research — May 2026.

The 7 papers the radar tracked in May 2026. The current radar lives on the research page. Back to the radar →

May 28HF Daily Papers

Distilling LLM Feedback for Lean Theorem Proving

Feedback Distillation improves post-training of reasoning models by using self-distillation with token-level supervision and privileged feedback from language models, offering bet…

Read paper · arxiv.org → Science Method May 28, 2026
May 28HF Daily Papers

Codifying the Judge: Scalable Evaluation via Program Distillation

LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability. We address these with a simple,…

Read paper · arxiv.org → Evals Benchmark May 28, 2026
May 18HF Daily Papers

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows actually prepare training data end to end.…

Read paper · arxiv.org → Evals Benchmark May 18, 2026
May 15HF Daily Papers

TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward

Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images more faithful to complex compositional prompts. We present…

Read paper · arxiv.org → Multimodal Method May 15, 2026
May 6HF Daily Papers

Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs

An axiomatic evaluation framework reveals systematic failures in latent thought representations of LLMs across multiple reasoning tasks, demonstrating that current representations fail to satisfy fundamental functional axioms consistently…

Read paper · arxiv.org → Evals Benchmark May 6, 2026
May 6HF Daily Papers

Token Time Continuous Diffusion for Language Modeling

In this paper we introduce token time continuous diffusion (TTCD), a new diffusion language model which (a) operates in continuous space, deterministically mapping Gaussian noise to a final token canvas with no further sampling, and…

Read paper · arxiv.org → Models Method May 6, 2026
May 6HF Daily Papers

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL

Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties become ineffective signals as model performance improves,…

Read paper · arxiv.org → Agents Method May 6, 2026