跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 5 天前AI 评分36

扩展与蒸馏文本嵌入以提升扩散语言模型可扩散性

AI 导读

研究发现,将嵌入模型从 T5 扩展到 T5Gemma-1、T5Gemma-2 可显著提升扩散语言模型生成性能,但原始 T5Gemma-2 嵌入区分度过高,导致连续扩散易落入无效嵌入。

正文

View PDF HTML (experimental)

Abstract:Diffusion language models (DLMs) offer a promising alternative to autoregressive (AR) language generation. Recent advances in continuous DLMs, which apply latent diffusion to continuous text embeddings, raise a practical question: which embedding makes the best latent space, i.e., the most diffusible? To answer this, we search through different embeddings and find that scaling the embedding model to stronger ones within the same family (T5 to T5Gemma-1 to T5Gemma-2) greatly improves generative performance. But the raw T5Gemma-2 embeddings are still not optimal. They are so discriminative that even the embeddings of plausible alternative words are separated, which makes the generation vulnerable to imperfect sampling. Consequently, continuous diffusion often fails to reach any of them and ends up at an invalid embedding instead. To address this, we distill T5Gemma-2 into a student encoder that learns the teacher's decoded probabilities as soft labels. Learning from such soft labels makes the student pull the alternative embeddings closer while maintaining the encoding-decoding mechanism. The distilled embeddings form a more connected and diffusible latent space, improving over the vanilla T5Gemma-2 embeddings. As a result, our medium-sized DLM achieves Gen. PPL 17.8 (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M on Gen. PPL.
Comments: 28 pages, 12 figures. Code is available at this https URL
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.01016 [cs.CL]
  (or arXiv:2610.01016v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.01016

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zekai Zhang [view email]
[v1] Thu, 1 Oct 2026 03:59:36 UTC (444 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org