跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 10 天前AI 评分41

BiasReducer:面向奖励模型的自适应偏见缓解框架

AI 导读

BiasReducer 提出一种轻量框架,仅编辑奖励模型的线性奖励头,并为每个新数据集自动选择相关编辑,无需重训练。它先用 SAE 风格编码器识别奖励模型敏感的属性(如长度、置信度),再学习调整奖励头的方向和幅度,最后按属性影响力排序筛选编辑。

正文

View PDF HTML (experimental)

Abstract:Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods either retrain the reward model or apply a fixed correction to one known bias, such as a preference for longer responses. Retraining requires additional data and computational resources, while existing editing methods require the target bias to be specified in advance and use a fixed edit for that bias. To this end, we propose BiasReducer, a lightweight framework that edits only the linear reward head and selects the relevant edits for each new dataset. First, BiasReducer uses a sparse autoencoder (SAE)-style encoder to learn which attributes (e.g., length and confidence) the reward model is sensitive to. Second, it learns how to reduce the reward model's dependence on each attribute by determining which direction to adjust the reward head and how much to adjust it. Third, for a new dataset, it ranks the attributes by their influence on reward scores, selects the relevant ones, and edits the reward model accordingly. BiasReducer consistently improves reward-model robustness to biases toward superficial response attributes. Across five reward models, BiasReducer-M improves the three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, outperforming the two training-based baselines. The gains transfer downstream, reducing unnecessary verbosity and sycophancy while maintaining comparable judged quality.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.32720 [cs.LG]
  (or arXiv:2609.32720v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.32720

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Shuang Liu [view email]
[v1] Sat, 26 Sep 2026 15:37:43 UTC (616 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org