HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分40
看清、说出、分类:LLM 涌现性失准的机制诊断与参数空间缓解
AI 导读
研究对 LLM 涌现性失准(EM)做二阶几何动态分析,发现方向性 Hessian 曲率集中在语义枢轴 token 上,有害-安全差距扩大主要源于安全梯度重叠下降。据此提出的参数级几何缓解框架,在 Qwen2.5-14B-IT 上将自由生成 EM 抑制最多 80.0%,并在四个开源指令模型家族中的另外三个(3B–20B)上验证同一有害子空间控制冻结 EM 响应的条件支持。
正文
Abstract:Safety-aligned LLMs can exhibit emergent misalignment (EM): narrow domain adaptation unexpectedly triggers catastrophic safety failures across unrelated domains. Prior static analyses leave training dynamics unmapped, while existing defenses rely on heuristics that degrade utility. We present a dynamic, second-order geometric study of EM. Tracking training trajectories reveals that directional Hessian curvature concentrates sharply on semantic pivot tokens. Grassmannian projections show that, in most settings, harmful-safe gap widens mainly because safe-gradient overlap declines. Leveraging these insights, we introduce a parameter-level Geometric Mitigation Framework that orthogonally projects empirical harmful gradient subspace out of parameter updates. On Qwen2.5-14B-IT, our defense suppresses free-generation EM by up to 80.0%; across the other three of four open-weight instruction-based model families (3B--20B), where single-layer behavioral EM is already near zero, teacher-forced evaluation shows same harmful subspace controls the conditional support of frozen EM responses. Crucially, these diagnostics unmask the illusion of behavioral safety: the same subspace remains measurable and steerable in models where behavioral EM is near zero. Code: this https URL.
| Comments: | Preprint |
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.34970 [cs.LG] |
| (or arXiv:2609.34970v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34970 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ruizhe Li [view email]
[v1]
Mon, 28 Sep 2026 11:51:35 UTC (2,727 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org