跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-07-17精选AI 评分74

从预训练到后训练:理解推理能力的强化学习缩放规律

AI 导读

一项新研究以国际象棋为受控实验平台,系统探究了预训练与强化学习(RL)在推理任务中的交互关系。研究发现,给定RL计算量下的后训练性能可由预训练损失准确预测,且RL奖励曲线的斜率随预训练token数近似线性提升。在1B参数数学语言模型上,更长预训练的检查点在RL下性能更高、提升更快。

推荐理由

这篇论文提出一个将预训练与RL联系起来的缩放定律,用国际象棋当试验场,发现RL在不同难度问题上的效果很不一样,做训练的人可以据此重新想想算力怎么分。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are vast and uncontrolled, making it hard to attribute behaviors to pretraining versus RL, and systematic compute sweeps across both stages are prohibitively expensive. To address these challenges, we use chess as a controlled testbed for studying reasoning across the full pretraining-to-post-training pipeline. We follow the standard LLM training pipeline by pretraining language models from 5M to 1B parameters on human chess games, supervised fine-tuning on synthetic reasoning traces, and running RL on chess puzzles with verifiable rewards. Using this framework, we find that the post-RL performance at given RL compute level is well-predicted from the pretraining loss, and slope of the RL reward curves improves approximately linearly with the pretraining tokens. Beyond scaling, we find that RL does not simply sharpen the SFT policy: on easy puzzles it amplifies correct moves the SFT policy already preferred, while on hard puzzles it surfaces correct moves that were nearly absent under SFT. We further test whether our findings transfer beyond chess by training a 1B language model on math-domain text, where the same predictive pattern emerges: longer-pretrained checkpoints reach higher post-RL performance and improve faster under RL. In sum, we provide a quantitative account of the pretraining-to-RL interface and a controlled testbed for studying the science of reasoning across the full pretraining-to-post-training pipeline.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2607.16097 [cs.LG]
  (or arXiv:2607.16097v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2607.16097

arXiv-issued DOI via DataCite

Submission history

From: Jingyan Shen [view email]
[v1] Fri, 17 Jul 2026 16:31:58 UTC (2,676 KB)
[v2] Sun, 9 Aug 2026 03:34:06 UTC (2,685 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org