HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分42
Physis-Lang:以自演化语言作为视频世界模型的物理表征
AI 导读
Physis-Lang 提出用自演化语言作为视频世界模型的物理表征,统一数据筛选、模型训练与视频生成,并构建 PhysCapBench 将物理过程拆解为原子断言、以召回率和精确率评估描述质量。在 Wan 与 Cosmos 骨干上,该方法在四个物理视频基准上持续提升物理合理性;基于开源 Cosmos3-Nano 的增强模型超过领先的闭源 Veo 3.1。
正文
Authors:Liming Lu, Xianzheng Ma, Wenkun He, Guanqi Zhan, Yilin Zhao, Junyu Chen, Mengyao Xu, Jiaojiao Fan, Wenhang Ge, Yuchao Gu, Yunze Liu, Boyi Li, Zhen Dong, Victor Prisacariu, Ming-Yu Liu, Song Han, Han Cai
Abstract:Video world models are expected to predict how the physical world evolves, yet they often produce visually plausible videos that violate basic physical principles. Existing approaches commonly assume that natural language is insufficient to represent the physical knowledge required for reliable generation, and therefore introduce additional visual, latent, numerical, or planning-based signals. We revisit this assumption and introduce Physis-Lang, a self-evolving framework that treats physical language as a shared and optimizable representation across data curation, model training, and video generation. Physis-Lang represents physical processes through language that describes their relevant entities, causes, interactions, governing principles, temporal evolution, and effects. To improve this representation, we construct PhysCapBench, which decomposes physical processes into atomic assertions and evaluates captions using recall and precision. An agentic loop iteratively analyzes assertion-level errors and refines the instruction used to produce physical captions. Physis-Lang further converts model deficiencies into textual descriptions and uses language-guided retrieval to identify visually diverse videos that cover missing physical processes. Experiments on four widely used physical video benchmarks with Wan and Cosmos backbones demonstrate consistent improvements in physical plausibility. Notably, starting from open-source Cosmos3-Nano backbones, our Physis-Lang-enhanced models surpass the leading proprietary Veo 3.1 model.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.40358 [cs.CV] |
| (or arXiv:2609.40358v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.40358 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xianzheng Ma [view email]
[v1]
Wed, 30 Sep 2026 17:59:51 UTC (3,936 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org