Communication Profiling and Characterization of Deep-Learning Workloads on Clusters With High-Performance Interconnects

Communication Profiling and Characterization of Deep-Learning Workloads on Clusters With High-Performance Interconnects
复制标题

具有高性能互连的集群上深度学习工作负载的通信分析和表征

DOI:
10.1109/mm.2019.2949986
复制
发表时间:
2020
期刊:
影响因子:
3.6
通讯作者:
D. Panda
D. Panda
中科院分区:
计算机科学3区
文献类型:
--
作者:
A. A. Awan;Arpan Jain;Ching;H. Subramoni;D. Panda

文献摘要

被引文献

相似文献

具有GPU的异质高性能计算系统配备了高性能互连,例如Infiniband,Omni-Path,PCIE和NVLINK。但是,文献中几乎没有捕获这些互连对分布式深度学习(DL)的性能影响。在本文中,我们选择了分布式培训中间件Horovod,除了标准MPI Microbenchs之外,还使用Tensorflow和Pytorch分析和概述了各种DNN培训工作负载。我们使用各种系统,具有CPU,例如Intel Xeon和IBM Power9,GPU,例如Volta V100和各种互连来分析以下指标:1)与Horovod的张量融合的消息大小; 2)无张量融合的消息大小; 3)MPI/NCCL呼叫的数量; 4)每个MPI/NCCL呼叫所花费的时间。我们观察到了不同平台上的两次消息大小的极端性能变化。为了解决这个问题,我们为HOROVOD设计了一个留言计划,说明了明显更顺畅的Alleduce延迟概况,并报告了我们观察到端到端培训改进的情况。
Heterogeneous high-performance computing systems with GPUs are equipped with high-performance interconnects like InfiniBand, Omni-Path, PCIe, and NVLink. However, little exists in the literature that captures the performance impact of these interconnects on distributed deep learning (DL). In this article, we choose Horovod, a distributed training middleware, to analyze and profile various DNN training workloads using TensorFlow and PyTorch in addition to standard MPI microbenchmarks. We use a wide variety of systems with CPUs like Intel Xeon and IBM POWER9, GPUs like Volta V100, and various interconnects to analyze the following metrics: 1) message-size with Horovod's tensor-fusion; 2) message-size without tensor-fusion; 3) number of MPI/NCCL calls; and 4) time taken by each MPI/NCCL call. We observed extreme performance variations for non-power-of-two message sizes on different platforms. To address this, we design a message-padding scheme for Horovod, illustrate significantly smoother allreduce latency profiles, and report cases where we observed improvement for end-to-end training.