HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分36
SAKI:面向同策略蒸馏的最大耦合路由教师监督方法
AI 导读
SAKI(Supervision Allocation with KL-constrained Interpolation)将 KL 约束的教师引导 rollout 与最大耦合结合,按 token 级接受/纠正事件路由监督:接受位置保留采样 token 的反向 KL 监督,纠正位置直接监督教师最高概率 token。
正文
Abstract:On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling and reuses realized accept/correction events to route token-level supervision. Accepted positions retain sampled-token reverse-KL supervision, while correction positions receive direct supervision on the teacher's highest-probability token. Under maximal coupling, the correction probability is exactly TV(p_t, q_t), so the same trust-region radius controls rollout deviation and upper-bounds intervention and specialized-supervision frequency. We further implement an engine-resident speculative verifier that preserves the exact-q trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x. Across seven mathematical reasoning benchmarks, SAKI improves the matched teacher-guided baseline in Mean@8 and Pass@8 for both 1.7B and 0.6B students. Placement controls and fixed-prefix analysis further support correction-triggered routing as a conflict-adaptive supervision signal.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.36601 [cs.AI] |
| (or arXiv:2609.36601v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.36601 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhenlin Wei [view email]
[v1]
Tue, 29 Sep 2026 03:18:09 UTC (446 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org