跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分44

SoL-Refiner:一步精修实现 4K 高分辨率视频生成

AI 导读

SoL-Refiner 是一种单步视频精修器,可将低分辨率模型输出经一次去噪直接转为 4K 视频,其训练流程包含高分辨率持续训练、强化学习后训练与最终的单步蒸馏。

正文

Authors:Haozhe Liu, Tian Ye, Shuchen Xue, Yitong Li, Junsong Chen, Haopeng Li, Jincheng Yu, Duomin Wang, Ruihua Zhang, Lei Zhu, Song Han, Enze Xie

View PDF HTML (experimental)

Abstract:High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a second sampling bottleneck. We present SoL-Refiner, a one-step video refiner that transforms low-resolution model outputs into 4K videos with a single denoising step. Our three-stage recipe combines high-resolution continual training, reinforcement learning (RL) post-training, and a final one-step distillation. We introduce Refiner-Bench, a video refinement benchmark constructed from the outputs of different video generators, and use a shared-input protocol to compare refiners at approximately 2K output resolution. At 2K, the one-step SoL-Refiner outperforms all external refiners on the VBench and UniPercept averages, while at $3840\!\times\!2176$ it improves both metrics over the three-step LTX-2.3 Refiner. With the complete acceleration stack, SoL-Refiner achieves an $8.91\times$ speedup in refinement latency over the same baseline in our 2K latency setting.
Comments: 15 pages
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2609.37969 [cs.CV]
  (or arXiv:2609.37969v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.37969

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Tian Ye [view email]
[v1] Tue, 29 Sep 2026 16:37:29 UTC (12,442 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org