HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分40
LIFT:通过 On-Policy 自蒸馏实现大视角变化下的未来布局视频生成
AI 导读
研究团队提出 LIFT,一个统一的图像到视频生成框架,用末帧布局作为显式控制信号,让用户指定未来视角中出现什么内容及其位置。由于稀疏布局引导比逐帧密集布局更难学习,LIFT 引入 on-policy self-distillation(OPSD),将密集布局教师模型的控制能力迁移给末帧布局学生模型。
正文
Authors:Shengxiang Ji, Boyang Wang, Haiyang Xu, Bingnan Li, Yucheng Mao, Zeyuan Chen, Xiaojun Shan, Xiang Zhang, Gang Hua, Jianwen Xie, Zezhou Cheng, Zhuowen Tu
Abstract:We introduce LIFT, a unified image-to-video generation framework that complements camera control with Layout-In-FuTure control, enabling users to specify what should appear in a future view and where it should appear. This addresses a practical need in controllable video generation: given an initial image, users often care not only about how the camera moves, but also about what the scene should look like at key future moments, especially the final frame. Existing camera controls specify viewpoint trajectories, while text prompts provide only coarse semantic guidance; neither precisely determines the content and spatial layout of future views. This limitation becomes particularly pronounced under large viewpoint changes, where the camera reveals regions that are not visible in the first frame. LIFT therefore uses the last-frame layout as an explicit control signal for the desired future scene. Since learning from such sparse layout guidance is substantially more challenging than conditioning on dense per-frame layouts, we introduce on-policy self-distillation (OPSD) to transfer the control capability of a dense-layout teacher to a last-frame-layout student. We further curate LIFT-Vista, a dataset featuring large viewpoint changes with camera and temporally consistent layout annotations. Experiments show that LIFT improves video quality, future-layout controllability, and camera controllability over other methods.
| Comments: | Project Page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.38146 [cs.CV] |
| (or arXiv:2609.38146v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38146 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shengxiang Ji [view email]
[v1]
Tue, 29 Sep 2026 17:57:20 UTC (24,293 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org