Enhancing Collective Communication in MCM Accelerators for Deep Learning Training

Enhancing Collective Communication in MCM Accelerators for Deep Learning Training
复制标题

DOI:
10.1109/hpca57654.2024.00069
复制
发表时间:
2024-03
期刊:
2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Sabuj Laskar;Pranati Majhi;Sungkeun Kim;Farabi Mahmud;A. Muzahid;Eun Jung Kim
Sabuj Laskar;Pranati Majhi;Sungkeun Kim;Farabi Mahmud;A. Muzahid;Eun Jung Kim
中科院分区:
其他
文献类型:
--
作者:
Sabuj Laskar;Pranati Majhi;Sungkeun Kim;Farabi Mahmud;A. Muzahid;Eun Jung Kim

文献摘要

相似文献

随着深度学习(DL)模型的广泛采用,对深度学习加速器硬件的需求也在增加。最重要的是,深度学习模型的规模正变得越来越大。为了适应这些模型,多芯片模块(MCM)成为实现大规模深度学习加速器的有效方法。虽然mcm在深度学习推理方面显示出了有希望的结果,但其在深度学习训练方面的潜力在很大程度上仍未被探索。目前的方法不能充分利用MCM加速器网状互连网络中的可用链路。为了解决这个问题,我们提出了两种新的基于网格的MCM加速器AllReduce算法——RingBiOdd和Three Tree Overlap (TTO)。RingBiOdd是一种基于环的算法,它通过使用双向互连创建两个单向环来增强AllReduce的带宽。另一方面,TTO是一种基于树的算法,通过重叠数据块来提高AllReduce的性能。TTO构造了三个拓扑感知的不相交树,并并行运行AllReduce操作的不同步骤。我们提出了一个详细的设计和实施所提出的方法。我们在7个深度学习模型上的实验结果表明,RingBiOdd比单向Ring AllReduce和MultiTree分别减少了50%和8%的训练时间。此外,与最先进的MultiTree和Bidirectional Ring AllReduce相比,TTO的训练时间分别减少了33%和29%。
With the widespread adoption of Deep Learning (DL) models, the demand for DL accelerator hardware has risen. On top of that, DL models are becoming massive in size. To accommodate those models, multi-chip-module (MCM) emerges as an effective approach for implementing large-scale DL accelerators. While MCMs have shown promising results for DL inference, its potential for Deep Learning Training remains largely unexplored. Current approaches fail to fully utilize available links in a mesh interconnection network of an MCM accelerator. To address this issue, we propose two novel AllReduce algorithms for mesh-based MCM accelerators - RingBiOdd and Three Tree Overlap (TTO). RingBiOdd is a ring-based algorithm that enhances the bandwidth of AllReduce by creating two unidirectional rings using bidirectional interconnects. On the other hand, TTO is a tree-based algorithm that improves AllReduce performance by overlapping data chunks. TTO constructs three topology-aware disjoint trees and runs different steps of the AllReduce operation in parallel. We present a detailed design and implementation of the proposed approaches. Our experimental results over seven DL models indicate that RingBiOdd achieves 50% and 8% training time reduction over unidirectional Ring AllReduce and MultiTree. Furthermore, TTO demonstrates 33% and 29% training time reduction over state-ofthe-art MultiTree and Bidirectional Ring AllReduce, respectively.