跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分44

改进的分布扩散模型:DiT-XL/2 在 ImageNet-256² 上 4 步达 4.48 FID

AI 导读

改进的分布扩散模型(DDM)通过将粒子扩展推迟到 Transformer 后层,并引入基于动力学区间的时间相关评分规则调度,解决了多粒子训练开销随粒子数增长、评分规则超参数全局固定的问题。

正文

View PDF

Abstract:Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a \emph{distributional} denoiser trained via a scoring rule objective, learning a stochastic approximation to $p(x_1 \mid x_t)$ rather than its conditional mean. However, scaling DDMs to modern image-generation settings faces two obstacles: (i) multi-particle training incurs overhead that scales with the number of particles, (ii) DDMs use globally fixed scoring rule hyperparameters, forcing a single trade-off across sampling budgets. We mitigate these limitations by deferring particle expansion to late transformer layers, and the hyperparameter trade-off by introducing time-dependent scoring rule schedules informed by the dynamical regimes of~\citet{Biroli2024}. Combined with a DiT-based latent setup, these changes make DDM training practical on class-conditional ImageNet-$256^2$, achieving 4.48 FID at 4 steps and 2.38 at 50 steps with DiT-XL/2, from a single model trained from scratch in one stage, without a teacher, self-distillation or JVPs. The result is a stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation. Code and pre-trained models available at this https URL.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2609.37147 [cs.CV]
  (or arXiv:2609.37147v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.37147

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Tommaso Martorella [view email]
[v1] Tue, 29 Sep 2026 09:41:56 UTC (34,480 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org