HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分41
结构化残差连接对 Diffusion Transformer 为何重要
AI 导读
研究提出将 Diffusion Transformer 的残差连接从被动求和改为主动检索机制,通过结构化连接设计让每个 Transformer block 选择性关注早期层表示。该方法训练迭代次数最多减少 1.73 倍,额外参数不到 0.1%,将 REPA-XL/2 的无引导 FID 从 5.9 降至 4.34,使用 classifier-free guidance 时达到 1.39 FID。
正文
Abstract:Diffusion Transformers (DiTs) have established themselves as a scalable backbone for high-fidelity image synthesis. However, unlike U-Net based diffusion models that rely on rigid, hand-crafted skip connections, DiTs predominantly use a uniform residual stream that integrates all preceding layers as a monolithic state. In this work, we rethink residual connections in diffusion transformers and propose to transform them from passive summation into an active retrieval mechanism optimized for image denoising. First, we conduct a systematic analysis of DiT's internal representation, revealing a latent preference for early-layer feature reuse and symmetric layer guidance. Motivated by this, we introduce a structured connectivity design that explicitly integrates local residual connections with long-range pathways. Instead of static skip connections or dense all-layer routing, our method enables each transformer block to selectively ``attend'' to critical earlier representations, dynamically retrieving spatial and semantic cues through direct, differentiable cross-depth paths. Experiments show that our adaptive connectivity leads to faster convergence, with up to $1.73\times$ fewer training iterations, and significant gains in FID and visual quality with less than $0.1\%$ additional parameters, further improving a strong REPA-XL/2 model from $5.9$ to $4.34$ FID without guidance and reaching $1.39$ FID with classifier-free guidance. Our findings suggest that adaptive cross-layer connectivity is a critical yet underexplored factor in diffusion transformers, and that incorporating structured information pathways provides a simple and effective direction for improving scalable generative models.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.33203 [cs.CV] |
| (or arXiv:2609.33203v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.33203 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yuhe Liu [view email]
[v1]
Sun, 27 Sep 2026 04:42:36 UTC (969 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org