跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分43

PreviewDiff:用多模态评审引导扩散潜变量搜索

AI 导读

PreviewDiff 是一种免训练测试时搜索方法,把扩散采样从标量搜索变为对中间潜变量的多模态评审引导搜索:在选定的去噪检查点解码部分预览,由多模态评审模型打分并给出自然语言批评,据此在语义提示词编辑和局部重新加噪的潜变量延续上分支,再对分支打分并选择性推进。在图像和视频生成基准上,它持续优于同等预算的 Best-of-N 选择和强标量搜索基线;消融显示更早介入和更大搜索宽度收益最大。

正文

View PDF

Abstract:Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails. We introduce PreviewDiff, a training-free test-time search method that turns diffusion sampling from scalar search into a multimodal critic-guided search over intermediate latents. At selected denoising checkpoints, PreviewDiff decodes a partial preview, asks a multimodal judge to score and critique it, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations. These branches are then scored and selectively rolled forward, allowing verifier compute to guide generation while the sample is still editable. Across image and video generation benchmarks, PreviewDiff consistently improves over budget-matched Best-of-N selection and strong scalar-search baselines. Ablations show that earlier interventions and increased search width provide the largest gains, while deeper search and additional semantic variants offer complementary improvements. PreviewDiff demonstrates that multimodal feedback is most useful not only as a final verifier, but as an active controller inside the denoising process.
Comments: 24 pages, 11 figures, 3 tables
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2609.36199 [cs.CV]
  (or arXiv:2609.36199v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.36199

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Vighnesh Subramaniam [view email]
[v1] Mon, 28 Sep 2026 20:00:41 UTC (8,901 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org