HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分49
Marathoner:超长时程自主智能体模型
AI 导读
研究团队提出自主智能体模型 Marathoner,通过后训练流水线赋予基座模型超长时程执行能力,在 5 个超长时程任务基准上稳定超越基座模型,甚至超过强闭源模型。该模型可连续工作 10+ 小时、执行 1000+ 次工具调用,训练数据来自 GitHub 仓库中新增 1000+ 行代码的大型发布 PR,并采用多任务链式合成与 Later Stage Bonus Reward 奖励策略。
正文
Abstract:Humans naturally possess the ability to work persistently toward long-term goals. Given a challenging task, humans can continuously work for months or even years to accomplish a specific objective. In this paper, we propose Marathoner, an autonomous agentic model possessing the ability of ultra-long-horizon execution. Specifically, we propose a comprehensive post-training pipeline to instill this critical capability into base model. For Ultra-Long-Horizon Task Synthesis, we leverage major release PRs containing 1000+ lines of new code from diverse GitHub repositories as the primary source for synthesizing challenging task-level data. Additionally, we introduce Multi-Task Chaining, which chains multiple generated tasks into a single more challenging task, enabling the synthesis of tasks with frontier-level difficulty. For rejection sampling finetuning, we combine strong teacher model with diverse harnesses to generate trajectories on our synthesized tasks and conduct supervised finetuning on base model with rejection sampled trajectories. For reinforcement learning, cold-started model performs real-world execution through harnesses in independent sandboxes during rollout process, effectively facilitating the acquisition of genuine ultra-long-horizon execution capability. We further propose a novel reward strategy, Later Stage Bonus Reward, which explicitly encourages model to perform meaningful maneuvers during later stages of execution. Through extensive evaluation on 5 benchmarks containing ultra-long-horizon tasks, Marathoner achieves consistent and substantial performance improvements over base model and even surpasses performance of strong proprietary model. Further analysis shows that Marathoner can consistently work for 10+ hours and conduct 1000+ tool calls on highly challenging tasks.
| Comments: | 30 pages, 4 figures |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.34378 [cs.CV] |
| (or arXiv:2609.34378v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34378 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ruiyang Zhang [view email]
[v1]
Mon, 28 Sep 2026 05:54:06 UTC (453 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org