HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分41
像素扩散模型的对抗训练后训练研究:恢复高频细节并提升分布保真度
AI 导读
研究首次系统探索像素扩散模型的对抗后训练:在预训练模型上保留原有扩散或流匹配目标,对非高噪声时间步的预测输出加入对抗损失,架构与采样流程不变。在两个像素骨干上,该方法同时提升了分布保真度、覆盖度、提示词对齐和感知质量;频谱分析显示原模型系统性缺失自然图像高频成分,对抗后训练可将其恢复。在测试的潜空间扩散配置下,同样流程未带来可比增益,几乎不增加解码后的高频功率。
正文
Abstract:Pixel diffusion models generate RGB images directly, avoiding the bottleneck of an autoencoder, yet their outputs still systematically underrepresent fine-scale natural-image statistics. We show that adversarial learning provides an effective post-training correction for this deficiency. Starting from a pretrained model, we retain its original diffusion or flow-matching objective and add an adversarial loss to the predicted output at non-high-noise timesteps, leaving the model architecture and sampling procedure unchanged. To our knowledge, this is the first systematic study of adversarial post-training for pixel diffusion. Across two pixel backbones, the method jointly improves distribution fidelity, coverage, prompt alignment, and perceptual quality. We further investigate why it works. Frequency-band and power-law analyses show that the original models systematically underproduce natural-image high-frequency content, while adversarial post-training restores this missing spectral power. In contrast, perceptual loss also increases high-frequency content but sacrifices distribution fidelity and prompt alignment. Nearest-neighbor, recall, and matched no-GAN SFT controls further rule out memorization, mode dropping, and additional optimization as simple explanations. Finally, we examine the boundary of this effect. Under the tested latent diffusion configurations, the same procedure does not produce comparable joint gains and adds almost no decoded high-frequency power. These results identify direct output access to the image statistics being corrected as a key factor governing when adversarial post-training succeeds.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2609.38170 [cs.CV] |
| (or arXiv:2609.38170v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38170 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xin Lin [view email]
[v1]
Tue, 29 Sep 2026 17:59:41 UTC (38,208 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org