HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 7 天前AI 评分39
SplitMoE:用分角色稀疏架构打破视频扩散模型的均匀性陷阱
AI 导读
针对传统 token-wise MoE 因均匀专家使用正则化导致的"均匀性陷阱",SplitMoE 提出分角色稀疏架构,将专家池拆分为语义专家与通用专家,配合原型引导路由和拉推正则化,让 token 按语义属性自然聚类。在等效激活参数预算下,SplitMoE 在收敛速度、路由一致性和视频生成质量上均优于传统负载均衡 MoE。
正文
Abstract:Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and pull-push regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.
| Comments: | Accepted as a Spotlight paper at NeurIPS 2026. Project page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.38140 [cs.CV] |
| (or arXiv:2609.38140v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38140 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yu Xu [view email]
[v1]
Tue, 29 Sep 2026 17:55:19 UTC (7,110 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org