跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分45

通过位置选择性自蒸馏从语言反馈中训练 LLM 评判模型

AI 导读

研究提出基于熵位移的位置掩码方法,从自然语言反馈中训练 LLM 评判模型,通过保留熵位移分布低尾位置来区分"上下文锐化"与"上下文扩散"两种机制。实验显示,掩码高熵位移位置可提升分布外泛化能力,所得自蒸馏评判模型在主观任务子类上比 GRPO 等结果监督 RL 方法高出 2-9 个百分点,在客观任务上保持竞争力。

正文

View PDF HTML (experimental)

Abstract:We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich language feedback (e.g., preference rationales) that naturally accompanies preference labels. Self-Distillation (SD) is one natural way to use this language feedback: the same model, conditioned on this feedback, acts as a teacher providing dense, position-level supervision. However, not all positions carry equally useful signal. Using the per-position entropy shift between teacher and student, we identify two regimes: context sharpening, where the teacher concentrates probability on a particular feedback-aligned criterion expression, and context spreading, where the teacher distributes probability across multiple feedback-aligned alternatives. We interpret these patterns as follows: sharpening encourages memorization of a particular criterion expression, whereas spreading promotes semantic understanding by preserving these alternatives. Motivated by this asymmetry, we introduce position masking based on the entropy shift that retains the lower tail of the entropy-shift distribution. Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD. The resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2609.38792 [cs.CL]
  (or arXiv:2609.38792v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.38792

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ilgee Hong [view email]
[v1] Wed, 30 Sep 2026 02:22:28 UTC (1,217 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org