HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分42
SlideDP:跨多 GPU 扩展宿主内存驻留的 LLM 微调
AI 导读
SlideDP 是面向共享宿主多 GPU 系统的同步数据并行运行时,通过维护单一权威宿主状态、解耦通信路径与状态布局,并对参数传输、梯度聚合和 CPU 更新做流水线化,实现超出 GPU 显存的 LLM 全参数微调。
正文
Abstract:Host-resident layer streaming enables full-parameter LLM fine-tuning beyond GPU memory, but data-parallel ranks compete for shared host resources. Replicated transfers amplify traffic, while strong scaling can expose host work as computation windows shrink. We present SlideDP, a synchronous data-parallel runtime for shared-host multi-GPU systems. It maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks. An analytical step-time model characterizes resource bottlenecks and pipeline exposure; runtime measurements guide communication, chunking, and activation policies under a GPU memory budget. In matched-batch sweeps, SlideDP achieves geometric-mean throughput ratios of 1.46-2.64$\times$ over SlideFormer, MegaTrain, and ZeRO-Offload. On four H100s, SlideDP approaches GPU-resident FSDP2 throughput for Qwen3-14B at a smaller batch size. With a larger batch, it processes over 1M tokens per step and exceeds FSDP2's measured peak throughput by 11.2%. Separately, it supports 256K-token sequences for the same model and fine-tunes Qwen2.5-72B on four RTX 4090 GPUs. Project page: this https URL.
| Comments: | 14 pages, 15 figures, 4 tables |
| Subjects: | Distributed, Parallel, and Cluster Computing (cs.DC) |
| Cite as: | arXiv:2609.34162 [cs.DC] |
| (or arXiv:2609.34162v1 [cs.DC] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34162 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ruijia Yang [view email]
[v1]
Mon, 28 Sep 2026 02:39:42 UTC (683 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org