跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分34

AdviSD:通过定向多轮自蒸馏让小型顾问模型指导前沿 LLM

AI 导读

AdviSD 让小型可训练顾问模型用自然语言指导冻结的 LLM 执行器,方法将结果导向的强化学习与选择性自蒸馏结合,用反馈条件下的顾问副本筛选反思修正。用 Qwen3-8B 顾问指导 Gemini 和 Claude 时,AdviSD 在 BFCL-v3 上比 advisor-GRPO 高 4.2-6.4 个百分点,在 EnvScaler 上高 3.9-5.1 分。

正文

View PDF HTML (experimental)

Abstract:A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2609.38142 [cs.AI]
  (or arXiv:2609.38142v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.38142

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Rishabh Agrawal [view email]
[v1] Tue, 29 Sep 2026 17:55:40 UTC (7,961 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org