跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分39

WorldAttention:面向交互式视频世界模型的高效注意力架构

AI 导读

WorldAttention 是一种面向交互式视频世界模型的注意力架构,通过 Hybrid Sparse Attention(HSA)与分层 KV Cache(HKV)的软硬件协同设计,在保留完整历史上下文的同时控制显存占用。在 VBench-Long 和 InterVBench 上,其主体一致性得分分别达到 0.9472 和 0.9668,持续超越此前 SOTA 方法。

正文

View PDF

Abstract:Leveraging the paradigm of autoregressive diffusion, text-conditioned interactive video world models aim to simulate temporally coherent environments guided by textual instructions. While enabling low-latency, long-duration generation is pivotal for embodied AI and simulation-based planning, current frameworks primarily rely on sliding-window mechanisms to bound computational complexity. However, this approach inherently sacrifices historical context, undermining the long-range interactive capabilities. Conversely, maintaining a full-history cache remains computationally prohibitive and memory-intensive: the quadratic complexity of attention leads to excessive computational overhead, while the linear growth of the KV cache inevitably leads to GPU memory saturation. To overcome these limitations, we propose WorldAttention, a system-oriented attention architecture that achieves high efficiency through the co-design of specialized attention kernels and hierarchical KV cache management. First, we introduce Hybrid Sparse Attention (HSA), which integrates linear global attention supplemented with head-adaptive sparse attention. Additionally, we design a Hierarchical KV Cache (HKV) that organizes historical KV pairs into semantically indexed pages across multi-tier memory, enabling fine-grained retrieval and controlled GPU residency. These two designs are supported by tailored kernels to effectively translate their theoretical efficiency into real-world performance. Extensive experiments on VBench-Long and InterVBench demonstrate that WorldAttention consistently surpasses prior state-of-the-art methods, achieving subject consistency scores of 0.9472 on VBench-Long and 0.9668 on InterVBench, respectively.
Comments: Website: this https URL, Code: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2609.34606 [cs.CV]
  (or arXiv:2609.34606v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.34606

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jinyuan Mao [view email]
[v1] Mon, 28 Sep 2026 08:37:04 UTC (14,642 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org