HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 5 天前AI 评分40
Ego2Act:评估第一视角视频生成中的目标导向操作
AI 导读
Ego2Act 是一个面向目标导向的第一视角视频生成基准,包含来自 110 个真实世界任务的 2,640 段视频,覆盖不同物体杂乱程度与多步复杂度。该基准要求模型根据初始场景图像和高层目标,生成手部操作物体完成任务的逼真第一视角视频,并配套推出无参考评估流程 Ego2ActJudge。
正文
Authors:Patrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana, Qinrong Cui, Erland Hilman Fuadi, Zayd M. K. Zuhri, Nanda Ryaas Absar, Ahmed Elshabrawy, Wilfried Ariel Mulyawan, Shoubin Yu, Yue Zhang, Mohit Bansal, Alham Fikri Aji
Abstract:Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamically evolve when executing goal-directed actions. While evaluating these capabilities is crucial, existing benchmarks focus mainly on single short actions or step-by-step instructions. This leaves multi-step physical reasoning underexplored, especially in egocentric video generation that requires planning to simulate proper execution to accomplish high-level goals by carrying out multiple real-world manipulations. We introduce Ego2Act, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity. Given an initial scene image and a high-level goal, Ego2Act evaluates whether video generation models can produce realistic egocentric videos of a hand manipulating objects to carry out the task. To support scalable evaluation, we also introduce Ego2ActJudge, a reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines. Our findings reveal that models' generated simulations often skip or partially execute steps, leaving later steps missing dependent states, which leads to unfulfilled goal. Furthermore, models consistently fail at fine-grained physical dynamics, particularly during complex object manipulation and persistent world modeling. We hope Ego2Act provides a rigorous testbed for advancing video models toward physically plausible, goal-directed simulation.
| Comments: | Preprint. 51 pages, 19 figures, 23 tables. Code, dataset and project website linked in the paper |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| ACM classes: | I.2.10; I.4.8 |
| Cite as: | arXiv:2610.01092 [cs.CV] |
| (or arXiv:2610.01092v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01092 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Patrick Irawan [view email]
[v1]
Thu, 1 Oct 2026 05:35:29 UTC (19,583 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org