HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 2026-05-21精选AI 评分70
RiT:在表示空间中使用原生扩散变换器已足够
AI 导读
本研究探讨预训练表示空间在流匹配学习中的优势。比较像素、SD-VAE与DINOv2特征后发现,尽管像素与DINOv2的内在维度相近,但DINOv2在几何统计特性(如有效秩、协方差条件等)上表现更优,使回归过程更稳定。基于此,我们提出了表示图像变换器(RiT),它使用冻结的DINOv2特征,通过x-prediction目标训练一个原生扩散变换器。在ImageNet 256×256生成任务上,RiT性能优于参数量更多的DiT^DH-XL模型,且生成的常微分方程仅需少量步骤即可高效求解。
推荐理由
这篇论文没发明新架构,但通过剖析DINOv2特征的统计属性,证明简单结构在表示空间也能做出SOTA,对做图像生成的人来说是个省钱省参数的好思路。
正文
Abstract:Flow matching with $x$-prediction -- regressing the clean data point rather than the ambient velocity -- is known to exploit low-dimensional manifold structure effectively in pixel space \cite{li2025back}. We ask whether a pretrained representation space, while containing a low-dimensional data manifold of comparable intrinsic dimensionality, offers a distribution more favorable for flow-matching learning. Comparing pixel, SD-VAE, and DINOv2 features along four geometric axes, we find that pixel and DINOv2 share nearly identical intrinsic dimensionalities (both $\hat{d}\!\approx\!33$) yet DINOv2 exhibits $7.3\times$ higher effective rank, $35\times$ better covariance conditioning, $11.5\times$ lower excess kurtosis, and $1.7\times$ lower on-manifold interpolation error; SD-VAE latents are consistently intermediate, indicating that the advantage stems from representation-learning objectives rather than mere compression. These statistical properties render the flow-matching regression well-conditioned and remove the need for the specialized prediction heads or Riemannian transport used by prior DINOv2 diffusion methods. We propose the \emph{Representation Image Transformer} (RiT): a vanilla Diffusion Transformer trained by $x$-prediction on frozen DINOv2 features, augmented only by a dimension-aware noise schedule and joint \texttt{[CLS]}-patch modeling. On ImageNet $256{\times}256$, RiT attains FID 1.45 without guidance and 1.14 with classifier-free guidance, outperforming DiT$^\text{DH}$-XL with $19\%$ fewer parameters (676M vs.\ 839M). The resulting ODE is efficiently solvable at coarse discretizations: with classifier-free guidance, $5$ Heun steps already reach FID 2.0 and $10$ steps reach 1.25, without distillation or consistency training. Code at this https URL.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2605.21981 [cs.CV] |
| (or arXiv:2605.21981v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2605.21981 arXiv-issued DOI via DataCite |
Submission history
From: Le Zhang [view email]
[v1]
Thu, 21 May 2026 04:21:43 UTC (7,139 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org