MalleTrain: Deep Neural Networks Training on Unfillable Supercomputer Nodes

MalleTrain: Deep Neural Networks Training on Unfillable Supercomputer Nodes
复制标题

DOI:
10.1145/3629526.3645035
复制
发表时间:
2024-04
期刊:
Proceedings of the 15th ACM/SPEC International Conference on Performance Engineering
影响因子:
--
通讯作者:
Xiaolong Ma;Feng Yan;Lei Yang;Ian T. Foster;M. Papka;Zhengchun Liu;R. Kettimuthu
Xiaolong Ma;Feng Yan;Lei Yang;Ian T. Foster;M. Papka;Zhengchun Liu;R. Kettimuthu
中科院分区:
其他
文献类型:
--
作者:
Xiaolong Ma;Feng Yan;Lei Yang;Ian T. Foster;M. Papka;Zhengchun Liu;R. Kettimuthu

文献摘要

相似文献

先来先服务调度可能会导致超级计算机上出现大量(高达 10%)暂时空闲的节点。由于 DNN 训练任务的灵活性,Liu 等人认识到此类未填充节点非常适合深度神经网络 (DNN) 训练。提出重新调整 DNN 训练任务以适应计划中的间隙,将其表述为混合整数线性规划 (MILP) 问题,并通过模拟证明了该方法的潜在优势。在这里,我们介绍 MalleTrain,这是一个系统,它提供了这种方法的第一个实际实现,并且通过允许它甚至可以用于运行前模型信息未知的 DNN 训练应用程序来进一步推广它。后一项创新的关键是使用轻量级在线作业分析顾问 (JPA) 来收集 DNN 作业的关键可扩展性信息,然后使用这些信息来实时动态优化资源分配。我们描述了 MalleTrain 架构,并展示了对超级计算机 GPU 集群和几个代表性 DNN 训练工作负载(包括神经架构搜索和超参数优化)进行详细实验评估的结果。我们的结果不仅证实了利用空闲超级计算机节点进行 DNN 训练的实际可行性,而且比之前的结果有了显着改进,将训练吞吐量提高了 22.3%,而无需用户提供作业可扩展性信息。
First-come first-serve scheduling can result in substantial (up to 10%) of transiently idle nodes on supercomputers. Recognizing that such unfilled nodes are well-suited for deep neural network (DNN) training, due to the flexible nature of DNN training tasks, Liu et al. proposed that the re-scaling DNN training tasks to fit gaps in schedules be formulated as a mixed-integer linear programming (MILP) problem, and demonstrated via simulation the potential benefits of the approach. Here, we introduce MalleTrain, a system that provides the first practical implementation of this approach and that furthermore generalizes it by allowing it to be used even for DNN training applications for which model information is unknown before runtime. Key to this latter innovation is the use of a lightweight online job profiling advisor (JPA) to collect critical scalability information for DNN jobs---information that it then employs to optimize resource allocations dynamically, in real time. We describe the MalleTrain architecture and present the results of a detailed experimental evaluation on a supercomputer GPU cluster and several representative DNN training workloads, including neural architecture search and hyperparameter optimization. Our results not only confirm the practical feasibility of leveraging idle supercomputer nodes for DNN training but improve significantly on prior results, improving training throughput by up to 22.3% without requiring users to provide job scalability information.