跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 6 天前AI 评分41

DC-SAE:面向更快扩散收敛的深度压缩语义自编码器

AI 导读

DC-SAE 是一种解耦紧凑语义自编码器,通过语义编码器实现高压缩比、并用像素级编码器保留细节,从而兼顾高压缩与快速扩散训练收敛。在 ImageNet 512×512 上实现 32 倍空间压缩,PSNR 29.79、gFID 3.37,较 DC-AE 分别提升 13.5% 和 54.9%。

正文

View PDF HTML (experimental)

Abstract:High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff between reconstruction fidelity and generation efficiency: high compression image encoder always increases the learning difficulty of diffusion training, resulting in slow model convergence. Recent representation autoencoders speed up the diffusion training by improving the latent feature's expressive capability by replacing VAE encoders with pretrained semantic encoders, yet they are typically limited to moderate compression and lose pixel-level details necessary for faithful reconstruction. To achieve both high compression and fast diffusion training, we propose DC-SAE, a Decoupled Compact Semantic Autoencoder designed for high-compression image generation with accelerated diffusion model convergence. DC-SAE consists of two key components: (1) a macro-level architecture design that leverages semantic encoders to enable higher compression ratios, and (2) a pixel-level encoder that preserves low-level details, ensuring high-fidelity image reconstruction. We empirically demonstrate that DC-SAE performs strongly on image generation tasks, achieving both compact latent representations and efficient training dynamics. Specifically, on the ImageNet dataset with $512 \times 512$ resolution, DC-SAE achieves $32\times$ spatial compression, with 29.79 PSNR and 3.37 gFID, substantially outperforming the previous state-of-the-art high-compression tokenizer baselines DC-AE by 13.5% and 54.9% on PSNR and gFID, respectively, maintaining comparable throughput and faster diffusion model training convergence. Beyond class-conditional generation, a $1.6$B-parameter DiT using DC-SAE achieves 0.84 on GenEval and 86.007 on DPG-Bench for text-to-image generation at $1024\times1024$ resolution.
Comments: 16 pages, 5 figures. Project page: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2609.39222 [cs.CV]
  (or arXiv:2609.39222v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2609.39222

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xu Huang [view email]
[v1] Wed, 30 Sep 2026 08:00:20 UTC (4,738 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org