跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分37

StoryEngine:面向视频叙事的基于状态的智能体框架

AI 导读

StoryEngine 是一个基于状态的智能体框架,通过分离权威语义规划与不可靠视觉观测来生成连贯的长篇视频故事。它维护实体位置与故事相关状态的结构化表示,传播事件引发的状态变化以定义每个镜头的起止状态,并构建可复用实体的规范参考,将状态与视觉约束编译为可执行的渲染计划。实验显示 StoryEngine 在叙事质量、叙事连贯性和视觉一致性等所有评估维度上均持续优于现有最先进方法。

正文

View PDF HTML (experimental)

Abstract:Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of story events and maintaining the video world state across shots. As a result, missing visual details may be reconstructed inaccurately, while visual drift may propagate across subsequent shots, undermining both narrative coherence and visual consistency. To address these challenges, we propose StoryEngine, a state-grounded agentic framework for video storytelling. StoryEngine establishes a separation between authoritative semantic plans and unreliable visual observations. Specifically, StoryEngine maintains a structured representation of entity placement and story-relevant states, and propagates event-induced changes to define the intended start and end states of each shot. To visually realize these states, StoryEngine constructs canonical references for recurring entities and environments, and compiles state and visual constraints into executable render plans. Meanwhile, to realize these states correctly, a bounded evaluation-guided repair loop further corrects local state inconsistencies. Together, these mechanisms preserve causal story progression and prevent local visual errors from propagating across shots. To comprehensively evaluate long-form storytelling, we construct a benchmark across diverse scenarios and visual styles, with metrics assessing storytelling quality, narrative coherence, and visual consistency. Experimental results demonstrate that StoryEngine consistently outperforms state-of-the-art methods across all evaluation dimensions, validating its effectiveness for coherent and consistent video storytelling.
Comments: 23 pages, 4 figures, submit to ICLR
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.33627 [cs.CV]
  (or arXiv:2609.33627v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.33627

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yingrui Wang [view email]
[v1] Sun, 27 Sep 2026 14:42:16 UTC (4,155 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org