跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分44

SkillGym:用自动可验证环境生成训练技能使用智能体

AI 导读

SkillGym 提出一套自动流水线,从互联网抓取技能并筛选可离线复现的工作流,通过 builder-reviewer 流程构建四类难度可控任务,每个任务配有参考解和可执行验证器,共建成 6.8k 个环境、收集 19k 条验证成功的轨迹用于监督微调。

正文

View PDF HTML (experimental)

Abstract:Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this work, we propose SkillGym, an automatic pipeline to build verifiable environments, collect trajectories, and train skill-use agents. SkillGym first crawls a large volume of skills from the internet, then keeps those whose workflows can run reproducibly offline. A builder-reviewer pipeline is used to construct difficulty-controlled tasks, spanning four task types, each with a reference solution and an executable verifier. With this pipeline, we build 6.8k environments and collect 19k verified successful trajectories for supervised finetuning. Finetuning on these trajectories improves LLMs of different families and sizes, from 2B to 122B parameters across four skill-use benchmarks; Our Qwen3.5-9B SFT model outperforms the 397B untrained model on two of them. Further analysis shows that training teaches agents to invoke skills, raising the rate of reading the relevant skill from 28% to 96%, and that the gains hold across reasoning structures, extending to task types that form a minority of the training data and to skills held out from training
Subjects: Artificial Intelligence (cs.AI)
ACM classes: I.2.7
Cite as: arXiv:2609.37539 [cs.AI]
  (or arXiv:2609.37539v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.37539

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Renxi Wang [view email]
[v1] Tue, 29 Sep 2026 13:21:38 UTC (357 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org