跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分45

模型内部表征何时有用?表征工程在 LLM 安全中的角色评估

AI 导读

一项匹配评估对比了表征工程与行为安全防护在 LLM 安全控制与监控上的表现。安全控制方面,DPO 整体控制最强且随训练数据增加而提升,但后续良性微调会削弱其安全性;表征引导仅在低数据场景(尤其高质量对比数据)保持竞争力。安全监控方面,专用文本监控器检测准确率最高,表征探针则以更低边际成本保持竞争力,监控引导干预还能在几乎不增加过度拒答的情况下恢复 DPO 因良性微调损失的大部分安全性。

正文

View PDF HTML (experimental)

Abstract:Reliable AI safeguards require both control mechanisms that reduce unsafe behavior and monitoring mechanisms that detect safety risks during model interactions. Established behavioral safeguards include alignment methods that optimize model outputs and text monitors that assess interaction text. Representation engineering instead reads or modifies internal model states, but the relative strengths of these approaches remain unclear because they are often evaluated under different settings. We present a matched evaluation across two tracks. For safety control, we compare DPO, a behavioral alignment method, with three representation steering methods across robustness, practicality, and granularity. DPO provides the strongest overall control and generally improves with increasing training data, although its safety can degrade after subsequent benign fine-tuning. Representation steering remains competitive primarily in low-data settings, particularly with high-quality contrastive data. For safety monitoring, we compare representation probes with fine-tuned and open-weight text monitors across full-response detection, early detection, and computational cost. Specialized text monitors achieve the strongest overall detection accuracy, while representation probes remain competitive at substantially lower marginal cost. Finally, monitor-guided interventions recover much of the safety lost by DPO after benign fine-tuning, with little additional over-refusal. Overall, representation engineering does not generally replace behavioral safeguards, but offers practical advantages under specific conditions and can provide complementary safety benefits.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2609.34771 [cs.AI]
  (or arXiv:2609.34771v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.34771

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Tianyi Guan [view email]
[v1] Mon, 28 Sep 2026 09:50:08 UTC (307 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org