SALaR: Scalable and Adaptive Designs for Large Message Reduction Collectives

SALaR: Scalable and Adaptive Designs for Large Message Reduction Collectives
复制标题

SALaR:大型消息缩减集体的可扩展和自适应设计

DOI:
--
复制
发表时间:
2018
期刊:
IEEE International Conference on Cluster Computing
影响因子:
--
通讯作者:
D. Panda
D. Panda
中科院分区:
--
文献类型:
--
作者:
Mohammadreza Bayatpour;J. Hashmi;Sourav Chakraborty;H. Subramoni;Pouya Kousha;D. Panda

文献摘要

被引文献

相似文献

到目前为止,消息传递界面(MPI)仍然是编程大规模科学应用程序的主要编程模型。 MPI的集体沟通运营非常重要,由于其沟通密集的性质和在科学应用中的使用。随着多/多核系统的出现以及深度学习应用的兴起,重新审视MPI集体,尤其是MPI Alleduce,以利用现代建筑提供的广泛的并行性。在本文中,我们采取了这一挑战,并提出了可扩展和适应性设计的大型消息集体(Salar)。由于它在深度学习框架中的使用,我们专注于MPI Alleduce,并提出了新设计,可以通过与高通量网络(例如Infiniband)一起利用现代多/多核的建筑特征来显着提高其性能。我们还提出了一个理论模型来分析通信和计算成本,并使用这些见解来指导我们的设计。对拟议的基于薪水设计的评估显示,对各种微基准和应用的最先进设计的性能取得了显着增长。
Message Passing Interface (MPI), thus far, has remained a dominant programming model to program large-scale scientific applications. Collective communication operations in MPI are of significant importance due to their communication intensive nature and use in scientific applications. With the emergence of multi-/many-core systems and rise of deep learning applications, it is important to revisit MPI collectives, particularly MPI Allreduce to exploit vast parallelism offered by modern architectures. In this paper, we take up this challenge and propose Scalable and Adaptive designs for Large message Reduction collectives (SALaR). We focus on MPI Allreduce due to its use in deep learning frameworks and propose new designs that can significantly improve its performance by exploiting architectural features of modern multi-/many-cores in tandem with high-throughput network such as InfiniBand. We also propose a theoretical model to analyze communication and computation cost and use these insights to guide our designs. The evaluation of the proposed SALaR based designs shows significant performance gains over state-of-the-art designs on a wide variety of micro-benchmarks and applications.