跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分38

EAPO:面向 LLM 推理探索的熵引导信用分配方法

AI 导读

针对 RLVR 中细粒度信用分配依赖辅助模型或额外采样的问题,研究者提出熵引导信用分配方法 EAPO,将归一化策略熵与响应优势符号耦合,对成功响应中的高熵决策加强强化、对失败响应中的低熵决策加强惩罚,同时衰减不确定位置的惩罚以保留恢复机会。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning with verifiable rewards (RLVR) enhances reasoning in large language models (LLMs) through outcome-level feedback, yet recent approaches to finer-grained credit assignment often require auxiliary models, additional sampling, or privileged information. Although policy entropy provides a readily available signal, prioritizing uncertain positions under both reinforcement and penalization concentrates penalties where failed responses still retain alternatives for recovery, which can suppress opportunities for exploration. To address this, we introduce Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method that treats success and failure asymmetrically. Specifically, motivated by the observation that success under uncertainty is less repeatable while confident failures tend to recur, EAPO couples normalized policy entropy with the sign of the response advantage to reinforce surprising success and correct repeated failure. It assigns stronger reinforcement to high-entropy decisions in successful responses and stronger penalties to low-entropy decisions in failed responses, while attenuating penalties at uncertain positions to preserve opportunities for recovery. By redistributing the response advantage across tokens, EAPO derives token-level credit directly from existing rollout signals without additional supervision. We validate EAPO on a range of reasoning tasks across both base and reasoning backbones, demonstrating that it achieves the best overall performance. We further show that EAPO promotes more effective exploration, broadening problem coverage and generating more diverse candidate answers.
Comments: Project page : this https URL
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2609.33781 [cs.LG]
  (or arXiv:2609.33781v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.33781

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Woongyeong Yeo [view email]
[v1] Sun, 27 Sep 2026 17:23:21 UTC (252 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org