HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分31
Neighborhood OPSD:用邻近参数扰动改进数学推理模型的在线自蒸馏
AI 导读
研究者提出 Neighborhood OPSD(N-OPSD),通过局部参数扰动构建冻结专家池,在相同参考解上下文下提供互补的参考对齐修正,并用 MaxPeak 选锚点 token、分位数选择匹配专家,学生仍用 OPSD 的 clipped forward-KL 目标学习。
正文
Abstract:On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises student-sampled prefixes. Standard OPSD uses one fixed parameter setting at every state, but nearby settings may offer additional supervision. We find that local parameter perturbations reveal complementary reference-aligned corrections under the same reference context. Different experts supply these corrections at different reference positions. Their pool covers more such positions than the unperturbed privileged teacher. We introduce Neighborhood OPSD (N-OPSD) to turn these corrections into supervision at student-visited states. Offline, greedy selection builds a compact pool of frozen experts by rewarding filtered reference-token gains beyond the pool's current best at each position. The highest-peak expert need not provide the best training target. Online routing therefore separates the anchor direction from its level of support. MaxPeak selects the anchor token, and quantile selection chooses among experts whose top token matches it. The student learns from the chosen expert's full next-token distribution through the clipped forward-KL objective inherited from OPSD. We evaluate on AIME 2024, AIME 2025, and HMMT February 2025. Across three independent runs per method, Neighborhood OPSD improves the three-benchmark Average@12 over OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B, respectively. Student-prefix continuations support using the pool beyond the reference trajectories used for selection. Matched ablations support filtered reference-token gains as a selection criterion. Accounting for overlap within the pool and routing by state further improve student accuracy. Inference uses only the distilled student.
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.39687 [cs.CL] |
| (or arXiv:2609.39687v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.39687 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xincheng Wei [view email]
[v1]
Wed, 30 Sep 2026 13:14:39 UTC (565 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org