HuggingFace Daily Papers(社区热门论文)· HuggingFace Daily Papers(社区热门论文)·· 8 天前AI 评分30
Nereus:面向 LLM 后训练的自适应并行运行时
AI 导读
Nereus 是一个面向 LLM 强化学习后训练的成本感知运行时,可动态调整并行策略以适配资源、序列长度和内存压力的变化。在真实数据 trace 中,在线 TP/PP 自适应将平均 step 延迟降低 27.7%;在 1,024 GPU 的 1,000 步运行中,六次切换仅占总运行时间的 0.079%。
正文
Abstract:Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a consequence, an execution plan that was initially suitable can then become slow or even infeasible over time. However, adapting a job whose models share GPUs entails significant challenges: deciding whether a new plan is worth the transition cost, reusing the job's distributed state, and coordinating GPU transfers across models and stages.
Nereus targets these challenges as a cost-aware runtime that adapts RL post-training jobs into efficient execution plans. Its low-overhead controller selects a memory-feasible global plan and admits the transition using a cost model calibrated against the running job. To estimate and execute a transition, Nereus represents the distributed state of each replica of a model-stage (one model in one stage) as an Elastic Model Unit. It then employs a global transition graph to order the transformations and GPU transfers of these units. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling. In a 1,000-step run reaching 1,024 GPUs, six transitions consume 0.079% of total run time. Nereus improves end-to-end 8B PPO throughput by 2.14--7.27$\times$ over OpenRLHF and by 1.10--1.47$\times$ over Verl across diverse clusters.
| Subjects: | Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI) |
| ACM classes: | C.2.4; C.1.4; I.2.6 |
| Cite as: | arXiv:2609.34645 [cs.DC] |
| (or arXiv:2609.34645v1 [cs.DC] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34645 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Songlin Jiang [view email]
[v1]
Mon, 28 Sep 2026 08:55:05 UTC (829 KB)
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org