跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分33

掩码扩散语言模型的状态自适应:何时真正有效?

AI 导读

研究将掩码扩散语言模型(MDM)的推理拆分为 score、cardinality、region、commitment、planning 五个维度,用"策略反转"衡量自适应机会,即最优候选动作相对验证集选定固定动作的单步收益优势。

正文

View PDF HTML (experimental)

Abstract:Masked diffusion language models (MDMs) admit flexible generation orders, making the unmasking strategy an inference decision. Existing methods vary in how they prioritize positions, control parallelism, restrict selection regions, revise predictions, or plan future denoising, yet it remains unclear when these choices should change during generation. We study this question through strategy reversals, where an alternative action becomes preferable to a fixed choice. We organize MDM inference into five axes--score, cardinality, region, commitment, and planning--and define adaptation opportunity as the one-step utility advantage of the best candidate action over a validation-selected fixed action. This view shows that adaptation value depends on both the frequency and magnitude of such reversals. Across three MDMs and ten tasks, adaptation opportunities are highly heterogeneous, with some regimes exhibiting concentrated and predictable one-step gains. This motivates selective adaptation: lightweight detectors calibrated on validation prompts identify high-opportunity states, capturing, for example, 56.9 percent of the candidate-set oracle opportunity by adapting only the top 10 percent of states on LLaDA-8B constrained JSON filling. Our transition-level results suggest that state adaptation is most useful when applied selectively rather than uniformly.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.33355 [cs.AI]
  (or arXiv:2609.33355v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.33355

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Injin Kong [view email]
[v1] Sun, 27 Sep 2026 08:27:32 UTC (280 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org