跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分29

E-MoE:面向非因子化扩散语言模型的增强型混合专家方法

AI 导读

研究提出 E-MoE,用 MoE 骨干的专家路由决策构建离散共享隐变量,将扩散语言模型的反向过程建模为因子化分布的混合,且不增加激活参数量。在合成多模态基准、二值化 MNIST 和 LM1B 上,E-MoE 相比因子化基线改善了少步生成质量。

正文

View PDF HTML (experimental)

Abstract:Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most. A recent line of work introduces a continuous Gaussian latent, trained as a variational autoencoder, to capture correlations across positions, but such approaches are prone to posterior collapse, where the latent is silently ignored. We propose Enhanced Mixture-of-Experts (E-MoE), which builds the reverse process as a mixture of factorized distributions over a discrete shared latent given by the expert-routing decisions of a Mixture-of-Experts (MoE) backbone, without increasing active parameters over the factorized baseline. Across synthetic multi-modal benchmarks, binarized MNIST, and LM1B, E-MoE improves few-step generation over factorized baselines.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2609.37533 [cs.CL]
  (or arXiv:2609.37533v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.37533

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Arseny Ivanov [view email]
[v1] Tue, 29 Sep 2026 13:18:50 UTC (2,677 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org