对齐如何路由:语言模型策略电路的定位、扩展与控制
研究定位了对齐训练语言模型的策略路由机制:中间层注意力门控检测内容并触发深层放大头输出拒绝信号。该机制在2B至72B参数的12个模型中普遍存在,随规模扩大从单头演变为跨层头带。实验证实门控层贡献不足1%输出却具因果必要性,单头消融在72B模型上效果减弱58倍。调节检测层信号可连续控制策略强度,甚至将拒绝转为有害回答。替换密码可使安全机制失效70-99%,而注入明文门控激活可恢复48%拒绝行为,表明安全能力仅被路由门控而非真正移除。
这篇论文定位了十二个模型中拒绝行为的电路机制,发现一个 gate-amplifier 路由结构,并且证明用简单密码就能完全绕过对齐,对安全审计和可解释性研究者是必读。
Abstract:We localize the policy routing mechanism in alignment-trained language models. An intermediate-layer attention gate reads detected content and triggers deeper amplifier heads that boost the signal toward refusal. In smaller models the gate and amplifier are single heads; at larger scale they become bands of heads across adjacent layers. The gate contributes under 1% of output DLA, yet interchange testing (p < 0.001) and knockout cascade confirm it is causally necessary. Interchange screening at n >= 120 detects the same motif in twelve models from six labs (2B to 72B), though specific heads differ by lab. Per-head ablation weakens up to 58x at 72B and misses gates that interchange identifies; at scale, interchange is the only reliable audit. Modulating the detection-layer signal continuously controls policy from hard refusal through evasion to factual answering. On safety prompts the same intervention turns refusal into harmful guidance, showing that the safety-trained capability is gated by routing, not removed. Thresholds vary by topic and by input language, and the circuit relocates across generations within a family even while behavioral benchmarks register no change. Routing is early-commitment: the gate fires at its own layer before deeper layers finish processing the input. An in-context substitution cipher collapses gate interchange necessity by 70 to 99% across three models, and the model switches to puzzle-solving rather than refusal. Injecting the plaintext gate activation into the cipher forward pass restores 48% of refusals in Phi-4-mini, localizing the bypass to the routing interface. A second method, cipher contrast analysis, uses plain/cipher DLA differences to map the full cipher-sensitive routing circuit in O(3n) forward passes. Any encoding that defeats detection-layer pattern matching bypasses the policy regardless of whether deeper layers reconstruct the content.
| Comments: | Code and data: this https URL. Accepted at the Mechanistic Interpretability Workshop at the 43rd International Conference on Machine Learning (ICML), 2026 |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2604.04385 [cs.CL] |
| (or arXiv:2604.04385v5 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2604.04385 arXiv-issued DOI via DataCite |
Submission history
From: Gregory Frank [view email]
[v1]
Mon, 6 Apr 2026 03:20:37 UTC (1,732 KB)
[v2]
Tue, 7 Apr 2026 12:41:36 UTC (1,834 KB)
[v3]
Mon, 13 Apr 2026 16:07:48 UTC (1,919 KB)
[v4]
Fri, 1 May 2026 15:41:40 UTC (1,919 KB)
[v5]
Mon, 29 Jun 2026 01:03:35 UTC (1,937 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org