HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分48
为何确定性 PRM 引导在离散扩散推理中表现不佳
AI 导读
研究显示,在同等前向计算预算下,确定性 PRM 引导的离散扩散语言模型推理不如独立采样加 ORM 重排。在 Dream-v0-Instruct-7B 上,GSM8K 每题 8 个候选时前者准确率 65.18%,后者达 75.13%,32 个候选时差距扩大到 12.69 个百分点。作者将其归因于引导阶段信号弱和 PRM 作为最终评判者能力不足,并发布了去噪状态语料与评测工具包。
正文
Abstract:Discrete diffusion language models (dLLMs) expose a denoised solution at every step, which makes process reward model (PRM) guidance look like a way to spend compute at test time. We show that once denoising, PRM scoring, and outcome reward model (ORM) scoring are charged in the same budget of forward passes, its deterministic form loses to a much simpler baseline. Our PRMs score intermediate denoising states and are trained on the correctness of the final answer. On Dream-v0-Instruct-7B with 8 candidates per GSM8K problem, keeping the candidate with the highest PRM score at every scoring step reaches 65.18%, while independent sampling plus an ORM reranker trained for the task reaches 75.13%. The gap grows to 12.69 percentage points (pp) with 32 candidates, and is 9.85 pp on MATH and 12.16 pp on MBPP. We trace it to two separable failures. First, guidance prunes on a weak signal: on GSM8K, PRM ROC-AUC falls from 0.77 to 0.54 as the mask ratio rises, a decay that persists when states are relabeled with fresh rollouts, and pruning lowers the best accuracy reachable from the candidate pool from 81.05% for independent samples to 67.30%. Second, on GSM8K and MATH, the PRM is a poor final judge: a sequential Monte Carlo sampler at the same budget restores that ceiling to 77.89%, yet selecting with the PRM gives 65.48%, on par with deterministic guidance, while a PRM retrained on final states matches the ORM on identical candidates. MBPP separates the two: there the PRM reaches 65.47% when reranking finished programs, on par with the ORM, but 50.88% when it guides denoising. The results point to two targets for dLLM guidance: keep correct partial solutions alive through early denoising, and leave the final choice to a verifier trained on final states. We release the corpus of denoising states with outcome labels and evaluation toolkit for reproducible comparisons at matched compute.
| Comments: | Accepted at NeurIPS 2026. 27 pages, 6 figures. Code: this https URL dataset and model: this https URL |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.35472 [cs.AI] |
| (or arXiv:2609.35472v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.35472 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yan Zhan [view email]
[v1]
Mon, 28 Sep 2026 15:54:05 UTC (1,627 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org