HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-07-22精选AI 评分70
RECAP:通过可解码性监督训练可验证的激活解释
AI 导读
研究发现自然语言自编码器(NLA)的重建分数无法验证逐声明的忠实性,模型可能依赖“私密代码”而非真实依据。作者提出RECAP方法,在目标模型上联合训练线性头以保持指定内容可解码。在Pythia-160M上,独立探针能可靠区分真假声明(AUC 0.96),并在对抗编辑下仍能标记谎言(AUC 0.95)。
推荐理由
这是近期可解释性领域最扎实的实证,它证明让模型自解释是危险的——重建测试会被内容和语言捷径欺骗,而训练时植入可解码性更可靠。虽仅在Pythia-160M上验证,但审计方法和思路对安全团队极为重要。
正文
Abstract:Natural-language autoencoders score explanations of hidden activations by reconstruction. An explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims. If flipping a claim does not change the reconstruction, the claim is never penalized. We show the test is passed in two ways, neither faithful. On a released Qwen-2.5-7B verbalizer, explanations reconstruct well above chance while ~2% of specific claims are ones the reconstruction depends on, so the score tracks gist, not specific facts. Under exact synthetic ground truth, standard training consistently develops co-adapted private codes (false wording the reconstruction depends on), and fixes that leave the target model unchanged do not help. We contribute two audit protocols, the comparison of grounding and truth and the swap to an independent evaluator, and RECAP (Readable Encodings via Co-trained Auxiliary Predictors), linear heads trained alongside the target model to keep designated content decodable. On RECAP-trained sandbox models, fresh verbalizers state the designated content truly and the codes vanish, at a +0.001-nat cost. This replicates on a pretrained Pythia-160M. The content becomes reliably decodable by a probe, though a fresh verbalizer conveys it only in part (truth 0.44-0.46 vs a near-zero control). For interpretability, high reconstruction does not certify individual claims. For AI safety, RECAP makes designated content checkable against a probe rather than asserted by prose a model can game. An independent probe ranks the verbalizer's true claims above its false ones (AUC 0.96 vs 0.82 without RECAP). Against an adversary that edits an explanation to maximize the score while lying (suppressing ~87% of its lie penalty), the RECAP probe still flags the lies (AUC 0.95) while the control probe collapses to chance (0.51).
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.20379 [cs.AI] |
| (or arXiv:2607.20379v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.20379 arXiv-issued DOI via DataCite |
Submission history
From: Hiskias Dingeto Dr [view email]
[v1]
Wed, 22 Jul 2026 17:10:23 UTC (325 KB)
[v2]
Wed, 19 Aug 2026 13:32:38 UTC (325 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org