跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分39

StructRL:面向长时程视觉-语言-动作任务的在线结构化强化学习

AI 导读

StructRL 是一个在线强化学习框架,通过将任务分解为可验证子任务、仅在前置子任务完成后给予中间奖励并按完成速度缩放奖励,为长时程 VLA 任务构建结构化中间监督。在 RoboCasa365 和 LIBERO-Long 上,配合 GR00T-N1.5 与 pi 0.5,StructRL 持续优于所评估的在线 RL 基线。结果表明可验证的结构化中间奖励能改善长时程 VLA 后训练,代码已开源。

正文

View PDF HTML (experimental)

Abstract:Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods provide reward only after the complete task succeeds. However, such terminal supervision is sparse and does not distinguish early failures from rollouts that make substantial partial progress. We propose StructRL, an online RL framework that constructs structured intermediate supervision from verifiable subtask completions. StructRL decomposes each task into verifiable subtasks, grants intermediate rewards only after the prerequisite subtasks have been completed, and scales each reward according to completion pace. Across RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, StructRL consistently outperforms evaluated online RL baselines. These results show that verifiable, structured intermediate rewards improve long-horizon VLA post-training. Code is available at this https URL.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Robotics (cs.RO)
Cite as: arXiv:2609.36352 [cs.CV]
  (or arXiv:2609.36352v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.36352

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sangmin Woo [view email]
[v1] Mon, 28 Sep 2026 22:34:43 UTC (1,827 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org