跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 9 天前AI 评分39

SMAT:简单高效的合并感知训练方法

AI 导读

研究者提出 SMAT(Simple MAT),通过将常见模型合并操作归纳为 Scale、Mask、Perturb 三种,联合优化专家损失与模拟合并参数下的期望损失,并引入周期性调度、算子融合和参数存储切换,使每步仅需一次前向和一次反向传播。在四个语言与视觉语言骨干模型上,SMAT 将五种合并方法的平均分较各骨干最强基线提升 1.07-2.16 分,训练时间开销较标准微调不足 2%。

正文

View PDF HTML (experimental)

Abstract:Model merging integrates the capabilities of multiple experts without joint retraining, but standard expert training optimizes task loss alone and does not guarantee good performance after merging. Merge-aware training (MAT) aims to improve merged performance, but existing methods do not fully account for common merging operations and add training cost. We observe that, from an expert's perspective, common merging methods can be described by three operations: Scale reweights its own update, Mask removes selected coordinates, and Perturb adds updates from other experts. Based on this view, we introduce SMAT (Simple MAT), which jointly optimizes expert loss and expected loss at simulated merged parameters generated by sampling scaling coefficients, masks, and additive noise. We further introduce periodic scheduling, kernel fusion, and parameter storage switching to make SMAT efficient, with one forward and one backward pass per step. Across four language and vision-language backbones, SMAT improves the mean score across five merging methods by 1.07-2.16 points over the strongest baseline for each backbone, with less than 2% training-time overhead over standard fine-tuning.
Comments: 19 pages, 6 figures, 7 tables. Code: this https URL
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as: arXiv:2609.33437 [cs.LG]
  (or arXiv:2609.33437v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.33437

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yanggan Gu [view email]
[v1] Sun, 27 Sep 2026 10:44:16 UTC (458 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org