跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-08-10精选AI 评分74

RynnValue:用时间距离扩展机器人价值基础模型

AI 导读

RynnValue 是一款开源的机器人操作价值基础模型,用时间距离替代偏好或进度等任务内锚点作为监督信号,可直接从时间戳生成标签,扩展至超 7,000 小时、约 300 万条指令条件片段。

推荐理由

用时间距离替代偏好标注,让价值模型能利用原始时间戳规模化训练,跨具身、任务和视角的泛化能力为无偏好数据的机器人策略提供了新的奖励接口。

正文

Authors:Dongchi Huang, Hongyin Zhang, Bohan Hou, Siteng Huang, Zhian Su, Hang Guo, Tong Lu, Jiahao Tang, Jianfei Yang, Donglin Wang, Peixi Peng, Mingxiu Chen, Deli Zhao, Xin Li

View PDF HTML (experimental)

Abstract:General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal anchors such as preferences or normalized progress, none of which transfer cleanly across embodiments and data sources. We introduce RynnValue, an open-source value foundation model for robotic manipulation that replaces these anchors with temporal distance, the directed cost-to-go from an observation to the language-specified goal. Because temporal-distance labels can be derived directly from timestamps, RynnValue scales to over 7,000 hours and roughly 3M instruction-conditioned clips without preference or progress annotations. To make temporal-value learning reliable at scale, we combine random temporal sampling, temporal-order shuffling, and value-isolation attention, suppressing shortcuts that would leave predictions insensitive to failures and regressions. Trained without preference labels, RynnValue attains an average Kendall's $\tau_a$ of 0.704 on RBM-EVAL-OOD, surpassing the fully preference-supervised state of the art (0.655) and more than doubling a progress-only counterpart (0.292), while generalizing zero-shot to unseen tasks, embodiments, and viewpoints. As a zero-shot reward model, RynnValue serves a range of downstream applications. Converted into dense rewards via potential-based shaping, it raises real-world policy success from 52.5% to 72.5% online and from 63.8% to 82.5% offline; used for data filtering, it improves multi-task behavior cloning success from 35.0% to 42.5%; and applied as inference-time value guidance, it lifts a frozen policy's success from 67.5% to 80.0%. These results establish temporal distance as a scalable supervision target and practical reward interface for generalist robot policies.
Comments: 32 pages, 7 figures
Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2608.09853 [cs.RO]
  (or arXiv:2608.09853v2 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2608.09853

arXiv-issued DOI via DataCite

Submission history

From: Dongchi Huang [view email]
[v1] Mon, 10 Aug 2026 17:09:37 UTC (4,319 KB)
[v2] Tue, 29 Sep 2026 13:33:11 UTC (6,217 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org