HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分44
ROSS:通过选择性监督从自生成 rollout 中再学习
AI 导读
ROSS 提出一种选择性监督方法,将历史轨迹完整保留为上下文,仅对选定的模型生成续写计算损失,从而复用自生成 rollout 中的行为经验。在 Qwen3.6-35B-A3B 上,ROSS 将六项基准的 MOPD 平均分从 58.40% 提升至 62.20%,SWE-bench Verified 从 64.20% 提升至 68.40%。
正文
Abstract:Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision. We introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), which preserves the full historical trajectory as context while applying loss only to selected model-generated continuations. Across domain-specific reinforcement learning, multi-teacher on-policy distillation, and agentic reinforcement learning, ROSS consistently improves upstream checkpoints and outperforms baselines across mathematics, code generation, instruction following, and software engineering. On Qwen3.6-35B-A3B, ROSS improves the six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%. These results show that self-rollout training leaves behind reusable behavioral experience that can yield further gains through offline supervised fine-tuning (SFT), without additional policy rollouts.
| Comments: | 21 pages, 6 figures, 10 tables |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.35954 [cs.LG] |
| (or arXiv:2609.35954v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.35954 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhiwei Zhang [view email]
[v1]
Mon, 28 Sep 2026 17:57:03 UTC (2,727 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org