跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分42

WorldLine:面向机器人操作的动作驱动视觉模拟器

AI 导读

WorldLine 是一个动作驱动的视觉模拟器,将可迁移的动力学学习与异构动作对齐解耦,基于超过 10,000 小时无动作机器人视频学习操作动力学,并用超过 2,000 小时、覆盖十余种本体的动作轨迹完成对齐。

正文

View PDF HTML (experimental)

Abstract:Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and coherent robot--object dynamics, while action-conditioned simulators depend on scarce, embodiment-specific data that are difficult to share across incompatible control spaces. We introduce WorldLine, an action-driven visual simulator that decouples transferable dynamics learning from heterogeneous action grounding. WorldLine learns manipulation dynamics from more than 10,000 hours of action-free robot videos and grounds them using over 2,000 hours of action trajectories across more than ten embodiments. An image-space action representation provides a shared control interface across embodiments, while multi-view and failure-enriched training with relational regularization improves interaction-sensitive prediction. Robot-focused few-step distillation enables efficient causal rollout while preserving action-critical motion. Across held-out and out-of-domain settings, WorldLine maintains strong visual quality and robot-motion agreement; on failed trajectories, it improves robot-mask IoU by 0.1626 over the strongest baseline. It predicts trajectory success with 74% mean accuracy across RoboTwin and AgiBot, one percentage point above the strongest baseline. Without RoboTwin training or adaptation, its rollouts improve task success by up to 21.4 percentage points over direct policy execution. Together, these capabilities make WorldLine a scalable and efficient visual simulator for policy evaluation and embodied planning. More results are available at \href{this https URL}{project page}.
Comments: A work about visual simulators for embodied AI
Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2609.38059 [cs.RO]
  (or arXiv:2609.38059v1 [cs.RO] for this version)
  https://doi.org/10.48550/arXiv.2609.38059

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Shenghe Zheng [view email]
[v1] Tue, 29 Sep 2026 17:26:22 UTC (47,371 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org