NV-group: link-efficient reduction for distributed deep learning on modern dense GPU systems

NV-group: link-efficient reduction for distributed deep learning on modern dense GPU systems
复制标题

DOI:
10.1145/3392717.3392771
复制
发表时间:
2020-06
期刊:
Proceedings of the 34th ACM International Conference on Supercomputing
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

被引文献

相似文献

NVIDIA NVLINK等高级织物正在实现密集图形处理单元(GPU)系统(例如DGX-2和Summit)的部署。通过广泛采用了用于分布式深度学习(DL)培训的大规模GPU系统,对于设计有效的沟通(例如Allreduce操作)至关重要,以便在大规模上实现近乎理想的加速。在本文中,我们通过NVLink-Aware合作的减少内核提出了一个链接效率的计划,以显着加速分布式深度学习应用程序的Allreduce操作。通过重叠计算和通信并最大化CPU和GPU之间所有可用的NVLink以及GPU之间的所有可用NVLink,我们证明了1,536 GPU的Allreduce的1.8倍性能提高了1,536 GPU,与最先进的GPU ARAT ARAT ARAT ARAT ARAT ARAT AREAD AREAD AREA ARE ARE ARE ARE AREA ARE AREADIA NVIDIA NCCL库。最后,我们分别在16-GPU DGX-2节点和Summit系统的192-GPU上展示了训练Resnet-50模型的93.9%和89.7%的缩放效率(即15倍和172x速度)。据我们所知,这是第一个为分布式DL培训而实现近乎理想的缩放效率的研究,并处理针对DGX-2和Summit簇等尖端系统量身定制的设计。
The advanced fabrics like NVIDIA NVLink are enabling the deployment of dense Graphics Processing Unit (GPU) systems such as DGX-2 and Summit. With the wide adoption of large-scale GPU-enabled systems for distributed deep learning (DL) training, it is vital to design efficient communication such as the Allreduce operation to achieve near-ideal speedup at scale. In this paper, we propose a link-efficient scheme through NVLink-aware cooperative reduction kernels to significantly accelerate Allreduce operations for distributed deep learning applications. By overlapping computation and communication and maximizing utilization of all available NVLinks between CPU and GPU, as well as among GPUs, we demonstrate 1.8X performance improvement of Allreduce on 1,536 GPUs compared to state-of-the-art GPU-Aware MPI and NVIDIA NCCL libraries. Finally, we demonstrate 93.9% and 89.7% scaling efficiency (i.e., 15X and 172X speedup) for training ResNet-50 models using TensorFlow on a 16-GPU DGX-2 node and on 192-GPUs of the Summit system, respectively. To the best of our knowledge, this is the first study that achieves near-ideal scaling efficiency for distributed DL training and deals with designs tailored for cutting-edge systems like DGX-2 and Summit clusters.