跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 10 天前AI 评分40

Allspark:用交替思维链实现弱到强迁移

AI 导读

Allspark 是一个通过交替思维链实现弱到强迁移的训练与推理框架:弱教师模型与同模型冻结副本交替推理、由冻结模型给出最终答案,推理时用更强的学生模型替换冻结伙伴,两者均保持固定。

正文

View PDF HTML (experimental)

Abstract:Recent progress in frontier models has renewed interest in large-scale reinforcement learning (RL), but the cost of generating large-model rollouts makes even testing RL recipes expensive. We ask whether reasoning improvements learned by a small, weak model can benefit a larger, stronger model without using the strong model's rollouts during training. We introduce Allspark, a training and inference framework for weak-to-strong transfer through alternating chains of thought. A weak teacher is trained alongside a frozen copy of the same model; the two alternate reasoning segments, and the frozen model produces the final answer. At inference time, a stronger student replaces the frozen training partner, while both models remain fixed. Because they communicate through text, the teacher can steer students from different model families and with different tokenizers. We study Allspark at two scales: controlled Qwen experiments across math and reasoning, and larger-scale Inkling experiments on ARC-AGI-2. The Inkling experiments show accuracy gains in within-family and cross-family settings, including transfer to Kimi and Nemotron, with benefits that vary across inference settings. These findings motivate reusing a trained weak teacher across strong students and examining the resulting accuracy--token tradeoff.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.32913 [cs.LG]
  (or arXiv:2609.32913v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.32913

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Kaizhao Liang [view email]
[v1] Sat, 26 Sep 2026 20:05:21 UTC (551 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org