跳到正文
原文
HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分42

SlideDP:跨多 GPU 扩展宿主内存驻留的 LLM 微调

AI 导读

SlideDP 是面向共享宿主多 GPU 系统的同步数据并行运行时,通过维护单一权威宿主状态、解耦通信路径与状态布局,并对参数传输、梯度聚合和 CPU 更新做流水线化,实现超出 GPU 显存的 LLM 全参数微调。

正文

View PDF HTML (experimental)

Abstract:Host-resident layer streaming enables full-parameter LLM fine-tuning beyond GPU memory, but data-parallel ranks compete for shared host resources. Replicated transfers amplify traffic, while strong scaling can expose host work as computation windows shrink. We present SlideDP, a synchronous data-parallel runtime for shared-host multi-GPU systems. It maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks. An analytical step-time model characterizes resource bottlenecks and pipeline exposure; runtime measurements guide communication, chunking, and activation policies under a GPU memory budget. In matched-batch sweeps, SlideDP achieves geometric-mean throughput ratios of 1.46-2.64$\times$ over SlideFormer, MegaTrain, and ZeRO-Offload. On four H100s, SlideDP approaches GPU-resident FSDP2 throughput for Qwen3-14B at a smaller batch size. With a larger batch, it processes over 1M tokens per step and exceeds FSDP2's measured peak throughput by 11.2%. Separately, it supports 256K-token sequences for the same model and fine-tunes Qwen2.5-72B on four RTX 4090 GPUs. Project page: this https URL.
Comments: 14 pages, 15 figures, 4 tables
Subjects: Distributed, Parallel, and Cluster Computing (cs.DC)
Cite as: arXiv:2609.34162 [cs.DC]
  (or arXiv:2609.34162v1 [cs.DC] for this version)
  https://doi.org/10.48550/arXiv.2609.34162

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ruijia Yang [view email]
[v1] Mon, 28 Sep 2026 02:39:42 UTC (683 KB)

来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org