Enabling Efficient Large-Scale Deep Learning Training with Cache Coherent Disaggregated Memory Systems

Enabling Efficient Large-Scale Deep Learning Training with Cache Coherent Disaggregated Memory Systems
复制标题

DOI:
10.1109/hpca53966.2022.00018
复制
发表时间:
2022-04
期刊:
2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Zixuan Wang;Joonseop Sim;Euicheol Lim;Jishen Zhao
Zixuan Wang;Joonseop Sim;Euicheol Lim;Jishen Zhao
中科院分区:
其他
文献类型:
--
作者:
Zixuan Wang;Joonseop Sim;Euicheol Lim;Jishen Zhao

文献摘要

被引文献

相似文献

现代深度学习(DL)训练是记忆的消费,受每个计算组件和跨设备通信带宽的记忆容量的约束。为了应对此类限制,当前的方法包括增加分布式培训中的并行性和优化设备间的通信。但是,模型参数通信已成为分布式DL培训中的关键性能瓶颈。为了提高参数通信性能,我们提出了粗糙的,这是分布式DL训练的分解内存扩展。粗糙是建立在现代高速缓存互连(CCI)协议和类似于MPI的集体通信上的基于同步的基础的,以便在Worker GPU中共享的训练数据和模型参数。为了在GPU和分解内存系统之间启用高带宽传输,我们提出了一个分散的参数通信方案,以将参数同步流量解除和本地化。此外,我们提出了动态张量路由和分区,以完全利用不同云计算系统各种的非均匀串行总线带宽。最后,我们设计了避免僵局和双重同步,以确保高性能参数同步。我们的评估表明,与最先进的MPI Alleduce通信相比,粗糙的DL训练速度高达48.3%。
Modern deep learning (DL) training is memory-consuming, constrained by the memory capacity of each computation component and cross-device communication bandwidth. In response to such constraints, current approaches include increasing parallelism in distributed training and optimizing inter-device communication. However, model parameter communication is becoming a key performance bottleneck in distributed DL training. To improve parameter communication performance, we propose COARSE, a disaggregated memory extension for distributed DL training. COARSE is built on modern cache-coherent interconnect (CCI) protocols and MPI-like collective communication for synchronization, to allow low-latency and parallel access to training data and model parameters shared among worker GPUs. To enable high bandwidth transfers between GPUs and the disaggregated memory system, we propose a decentralized parameter communication scheme to decouple and localize parameter synchronization traffic. Furthermore, we propose dynamic tensor routing and partitioning to fully utilize the non-uniform serial bus bandwidth varied across different cloud computing systems. Finally, we design a deadlock avoidance and dual synchronization to ensure high-performance parameter synchronization. Our evaluation shows that COARSE achieves up to 48.3% faster DL training compared to the state-of-the-art MPI AllReduce communication.