跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分35

ThinkV2V:释放 MLLM 推理能力,实现指令引导的视频编辑

AI 导读

ThinkV2V 是一个推理驱动的指令引导视频编辑框架,通过 MLLM-to-DiT 架构让 MLLM 在视觉生成前先对源视频和指令进行显式思考,生成精炼的条件信号。它结合渐进式课程训练与推理时思考扩展,并配套 ThinkV2V-150K 数据集和 ThinkV2V-Bench 基准。实验显示其 5B 规模 DiT 模型在复杂与标准编辑场景均达 SOTA,大幅超越 10B 规模基线。

正文

Authors:Donghao Zhou, Haoyang He, Fan Zhang, Hao Yang, Guisheng Liu, Xin Gao, Zhongwei Wan, Xingyuan Bu, Jie Wang, Qiangpeng Yang, Shilei Wen, Chi-Wing Fu, Pheng-Ann Heng

View PDF HTML (experimental)

Abstract:Instruction-guided video editing has made significant progress, yet existing methods use multimodal large language models (MLLMs) primarily as semantic encoders, so they often fall short in working with implicit edits that require causal or semantic reasoning. To bridge this fundamental gap in video editing, we propose ThinkV2V, a reasoning-driven framework for complex instruction-guided video editing, explicitly activating MLLM thinking before visual generation. At its core, ThinkV2V builds on a practical MLLM-to-DiT architecture to turn explicit thinking over the source video and instruction into refined conditioning signals for video editing. Further, we equip it with a dedicated training and inference recipe, combining Progressive Curriculum Training, which gradually cultivates the model from basic editing to reasoning-intensive cases, with Inference-Time Thinking Scaling, which iteratively refines candidate prompts and selects the most reliable one, to better elicit reasoning in challenging editing scenarios. We also curate the ThinkV2V-150K dataset and introduce ThinkV2V-Bench to support training and evaluation of video editing with implicit intent and causal reasoning. Experimental results demonstrate the state-of-the-art performance of ThinkV2V on both complex and standard editing scenarios, in which our 5B-scale DiT model substantially outperforms larger 10B-scale baselines.
Comments: Project page: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2609.38541 [cs.CV]
  (or arXiv:2609.38541v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.38541

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Donghao Zhou [view email]
[v1] Tue, 29 Sep 2026 20:58:42 UTC (23,560 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org