跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 5 天前AI 评分57

论文提出 Sharpening Tax 量化 RL 后训练的解覆盖损失

AI 导读

论文研究了 LLM 的 RL 后训练是否只锐化基座已有行为的问题,发现在智能体任务上,带轻量推理框架的预训练模型虽 pass@1 更低,但在足够测试预算下 pass@K 常超过后训练模型。

正文

View PDF HTML (experimental)

Abstract:An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.01509 [cs.AI]
  (or arXiv:2610.01509v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.01509

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Changdae Oh [view email]
[v1] Thu, 1 Oct 2026 11:46:00 UTC (7,182 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org