HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分44
蒸馏防御在强化学习后轻易失效
AI 导读
论文指出,现有蒸馏攻击防御通常只在蒸馏后立即评估,隐含假设攻击者不再继续训练,但更现实的威胁模型包含蒸馏后的强化学习。结果显示,强化学习会降低蒸馏攻击门槛,仅用当前 API 易得数据即可窃取闭源模型推理能力,效果相当于提取完整隐藏推理轨迹的复杂攻击。任何泄露足够信息以重建近似推理轨迹的蒸馏防御都可能失效。
正文
Abstract:Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e., "distill") their own models on these traces. Existing defenses against distillation attacks are typically evaluated immediately after distillation, implicitly assuming attackers do not train their models any further. In this paper, we argue that a more realistic threat model includes further training with reinforcement learning after distillation. A misspecified threat model can give a false sense of security -- some defenses that seem effective after distillation can be broken after subsequent reinforcement learning. Practically, reinforcement learning lowers the bar for a distillation attack to be effective. We show that simple attacks can steal reasoning capabilities from existing closed-source language models using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract the full hidden traces. Results indicate that any distillation defense that leaks sufficient information to reconstruct approximate reasoning traces is likely ineffective. We conclude by discussing broader implications and batch-level distillation defenses which could be more effective.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR) |
| Cite as: | arXiv:2609.35699 [cs.LG] |
| (or arXiv:2609.35699v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.35699 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yonatan Gideoni [view email]
[v1]
Mon, 28 Sep 2026 17:40:32 UTC (738 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org