跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分44

NEEDLE:通过权重正交化移除 LLM 后门

AI 导读

研究者提出免训练的后门移除方法 NEEDLE,通过激活向量估计后门方向与拒绝子空间,再施加顺序权重正交化,在压制后门的同时避免改变拒绝相关表征。该方法无需干净参考模型或原始投毒训练数据,在多种模型家族与攻击类型上取得最低平均攻击成功率,其中代码注入攻击为 0%,同时 KL 散度最低,能力与安全性变化极小。

正文

View PDF HTML (experimental)

Abstract:Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded model performance and safety. We propose NEEDLE, a training-free method for targeted backdoor removal. Once a trigger has been identified, our method estimates a backdoor direction and a refusal subspace through activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preventing changes in refusal-related representations. NEEDLE requires neither a clean reference model nor the original poisoned training data. Evaluation is conducted across multiple model families and attack types. NEEDLE achieves the lowest mean Attack Success Rate (ASR) among the evaluated defences, including 0% on challenging code injection attacks, while resulting in the lowest KL divergence and minimal changes in capability and safety.
Subjects: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
Cite as: arXiv:2610.00348 [cs.CR]
  (or arXiv:2610.00348v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.00348

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: George Drayson [view email]
[v1] Tue, 29 Sep 2026 20:35:24 UTC (271 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org