Communication Algorithm-Architecture Co-Design for Distributed Deep Learning

Communication Algorithm-Architecture Co-Design for Distributed Deep Learning
复制标题

DOI:
10.1109/isca52012.2021.00023
复制
发表时间:
2021-06
期刊:
2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Jiayi Huang-;Pritam Majumder;Sungkeun Kim;A. Muzahid;K. H. Yum;Eun Jung Kim
Jiayi Huang-;Pritam Majumder;Sungkeun Kim;A. Muzahid;K. H. Yum;Eun Jung Kim
中科院分区:
其他
文献类型:
--
作者:
Jiayi Huang-;Pritam Majumder;Sungkeun Kim;A. Muzahid;K. H. Yum;Eun Jung Kim

文献摘要

被引文献

相似文献

大规模分布式深度学习训练使得更复杂的深度神经网络模型能够从更大的数据集中学习复杂的任务。特别是,分布式随机梯度下降集中调用所有约简操作来进行梯度更新,这在迭代训练时期占主导地位。在这项工作中,我们确定了广泛使用的所有减少算法的效率低下,算法架构协同设计的机会。我们提出了多树全约简算法与拓扑和资源利用率的意识,有效和可扩展的全约简操作,适用于不同的互连拓扑结构。此外,我们共同设计的网络接口,调度和协调的所有减少无竞争通信的消息,协同工作的算法。简化了流量控制,充分利用了大梯度交换的大容量数据传输。我们使用不同的全简化数据大小来评估协同设计进行综合研究,证明其在各种互连网络拓扑结构上的有效性,以及用于真实的工作负载实验的最先进的深度神经网络。结果表明,MultiTree实现了2.3倍和1.56倍的通信加速,以及高达81%和30%的训练时间减少相比,环所有减少和国家的最先进的方法,分别。
Large-scale distributed deep learning training has enabled developments of more complex deep neural network models to learn from larger datasets for sophisticated tasks. In particular, distributed stochastic gradient descent intensively invokes all-reduce operations for gradient update, which dominates communication time during iterative training epochs. In this work, we identify the inefficiency in widely used all-reduce algorithms, and the opportunity of algorithm-architecture co-design. We propose MultiTree all-reduce algorithm with topology and resource utilization awareness for efficient and scalable all-reduce operations, which is applicable to different interconnect topologies. Moreover, we co-design the network interface to schedule and coordinate the all-reduce messages for contention-free communications, working in synergy with the algorithm. The flow control is also simplified to exploit the bulk data transfer of big gradient exchange. We evaluate the co-design using different all-reduce data sizes for synthetic study, demonstrating its effectiveness on various interconnection network topologies, in addition to state-of-the-art deep neural networks for real workload experiments. The results show that MultiTree achieves 2.3× and 1.56× communication speedup, as well as up to 81% and 30% training time reduction compared to ring all-reduce and state-of-the-art approaches, respectively.