HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分45
稀疏 Crosscoders 揭示 On-Policy Distillation 机制:学生模型只是重新加权已有特征
AI 导读
研究用稀疏 crosscoders 分析 LLM 推理中的 on-policy distillation(OPD),发现 OPD 既不创造新特征也不传递教师模型自身特征,学生模型超 98% 高频使用特征的激活率变化在 20% 以内。
正文
Abstract:On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher, yet what OPD actually distills into the student's internal representations remains unclear. We study this question with sparse crosscoders, which learn one feature dictionary shared by the student before and after OPD and the teacher. Standard crosscoder analyses, however, identify model-specific features but cannot tell how a model's use of its features changes, since all models are encoded into one set of feature activations. We therefore propose the swap readout, which reads each student checkpoint's feature activations on its own, measuring how training changes the student's use of each feature, even for checkpoints unseen by the crosscoder. Across three OPD settings, we find that OPD neither creates features nor passes on the teacher's own, and leaves the firing rates of over 98% of the student's frequently used features within 20%. We further examine the SFT warm-up on the teacher's rollouts that commonly precedes OPD and makes it more effective. Rather than adding features, the warm-up reweights the shared ones in two ways. First, it already raises and lowers many of the features that OPD later raises and lowers, doing part of OPD's work in advance. Second, it changes features that OPD alone would not, notably those for conversation format, reasoning style, and mathematical notation, and these changes persist through OPD. Imposing this reweighting on a directly distilled student's features, without changing its weights, brings its accuracy close to that of the warmed-up student, whereas the same change on shuffled features does not. Together, these findings suggest that OPD reweights existing features rather than acquiring new ones: the student learns from the teacher how to use the features they already share.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.35210 [cs.CL] |
| (or arXiv:2609.35210v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.35210 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zichao Yu [view email]
[v1]
Mon, 28 Sep 2026 14:03:08 UTC (800 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org