跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分42

EditWorld:面向可交互世界的精准编辑与灵活引用视频世界模型

AI 导读

EditWorld 是一个支持精准编辑与灵活引用的视频世界模型,可在自回归生成过程中流式接收编辑指令和参考图像,将世界建模从探索扩展到精准修改。它引入 Gated Causal Attention 处理时变编辑条件与参考图像,并用 Sparse Context 机制维持有界历史上下文以支持长时程推理。

正文

View PDF HTML (experimental)

Abstract:We present EditWorld, a video world model for precise editing and flexible referencing in interactable worlds. Existing video world models primarily focus on navigation, letting users explore generated worlds but offering limited control over how existing world content is modified. EditWorld extends world modeling from exploration to precise modification by streaming editing instructions and reference images during autoregressive generation. To support these capabilities, EditWorld introduces Gated Causal Attention for temporally varying editing conditions and reference images, together with a Sparse Context mechanism that maintains a bounded historical context for long-horizon inference. We further adopt joint autoregressive and bidirectional training with annealed self-resampling, and construct a dedicated data synthesis and annotation pipeline that provides supervision for world editing. We also present WBench-Editing to systematically evaluate streaming world editing capabilities. EditWorld achieves the best overall performance on WBench-Editing with an overall score of 73.8 and an editing score of 80.0, substantially outperforming existing methods on editing-related metrics. this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.34470 [cs.CV]
  (or arXiv:2609.34470v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.34470

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xinyao Liao [view email]
[v1] Mon, 28 Sep 2026 07:26:50 UTC (8,094 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org