Hy-Fi: Hybrid Five-Dimensional Parallel DNN Training on High-Performance GPU Clusters

Hy-Fi: Hybrid Five-Dimensional Parallel DNN Training on High-Performance GPU Clusters
复制标题

Hy-Fi:高性能 GPU 集群上的混合五维并行 DNN 训练

DOI:
10.1007/978-3-031-07312-0_6
复制
发表时间:
2022
期刊:
Proceedings International Conference on High Performance Computing
影响因子:
--
通讯作者:
Panda, DK.
Panda, DK.
中科院分区:
--
文献类型:
--
作者:
Jain, A;Shafi, A.;Anthony, Q.;Kousha, P.;Subramoni, H.;Panda, DK.

文献摘要

相似文献

高性能计算(HPC)的最新进展使深度学习(DL)模型能够通过利用多个处理器来实现最先进的性能。数据并行是一种在每个处理器上复制DL模型的策略,这对于像NVIDIA GPU上的AmoebaNet这样的模型是不可能的。层并行通过在每个GPU上放置一个或多个层来避免这种限制,但仍然无法在高分辨率图像上训练AmoebaNet等模型。我们建议Hy-Fi:Hybrid Five-Dimensional并行性;一种利用五个并行性维度(数据、模型、空间、管道和双向并行性)的系统,可实现核外模型和层的高效分布式训练。Hy-Fi还提出了通信级优化来整合这些维度。我们分别报告了层并行和流水线并行的加速比.我们在AmoebaNet和ResNet模型上演示了Hy-Fi在多达2048个GPU上的运行。此外,我们使用Hy-Fi来支持高分辨率图像的DNN训练,包括8,1928,192和16,38416,384。
Recent advances in High Performance Computing (HPC) enable Deep Learning (DL) models to achieve state-of-the-art performance by exploiting multiple processors. Data parallelism is a strategy that replicates the DL model on each processor, which is impossible for models like AmoebaNet on NVIDIA GPUs. Layer parallelism avoids this limitation by placing one or more layers on each GPU, but still cannot train models like AmoebaNet on high-resolution images. We propose Hy-Fi: Hybrid Five-Dimensional Parallelism; a system that takes advantage of five parallelism dimensions—data, model, spatial, pipeline, and bi-directional parallelism—which enables efficient distributed training of out-of-core models and layers. Hy-Fi also proposes communication-level optimizations to integrate these dimensions. We report up toandspeedups over layer and pipeline parallelism, respectively. We demonstrate Hy-Fi on up to 2, 048 GPUs on AmoebaNet and ResNet models. Further, we use Hy-Fi to enable DNN training on high-resolution images, including 8,1928,192 and 16,38416,384.