跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分39

自然语言自编码器如何挑选最具信息量的 token 位置

AI 导读

研究在 470 万条关于提示注入与隐瞒行为的解释上,比较了模型计算信号与仅用聊天结构训练的排序器,发现聊天结构通常能选出更相关的解释,且无需为位置选择做一次模型前向传播。在四个数据集中的三个上,只解释 5% 的位置就能保留解释全部位置时几乎全部的成功率。预训练 verbalizer 还能在无需额外训练的情况下还原模型经微调学会隐瞒的词语。

正文

View PDF HTML (experimental)

Abstract:Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across $4.7$ million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just $5\%$ of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.37040 [cs.CL]
  (or arXiv:2609.37040v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.37040

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Federico Torrielli [view email]
[v1] Tue, 29 Sep 2026 08:50:02 UTC (1,101 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org