Stash: A Comprehensive Stall-Centric Characterization of Public Cloud VMs for Distributed Deep Learning

Stash: A Comprehensive Stall-Centric Characterization of Public Cloud VMs for Distributed Deep Learning
复制标题

DOI:
10.1109/icdcs57875.2023.00023
复制
发表时间:
2023-07
期刊:
2023 IEEE 43rd International Conference on Distributed Computing Systems (ICDCS)
影响因子:
--
通讯作者:
Aakash Sharma;Vivek M. Bhasi;Sonali Singh;Rishabh Jain;Jashwant Raj Gunasekaran;Subrata Mitra;
Aakash Sharma;Vivek M. Bhasi;Sonali Singh;Rishabh Jain;Jashwant Raj Gunasekaran;Subrata Mitra;
中科院分区:
其他
文献类型:
--
作者:
Aakash Sharma;Vivek M. Bhasi;Sonali Singh;Rishabh Jain;Jashwant Raj Gunasekaran;Subrata Mitra;

文献摘要

相似文献

深度神经网络(DNN)因其解决图像识别、自动驾驶和自然语言处理等复杂问题的能力而越来越受欢迎。它们日益复杂,再加上使用大量训练数据(以达到可接受的精度),因此需要使用 GPU 和其他加速器。此类加速器通常价格昂贵,用户必须支付高昂的前期成本才能获得它们。对于不经常使用的情况,用户可以利用公共云来降低高昂的采购成本。然而,由于公有云中可用的硬件实例(特别是 GPU 实例)多种多样,用户从成本/性能的角度做出适当的选择变得具有挑战性。在这项工作中,我们尝试通过以下方式解决这个问题:(i) 引入全面的分布式深度学习 (DDL) 分析器 Stash,它可以确定 DDL 遇到的各种执行停顿;(ii) 使用 Stash 通过在各种公共云 GPU 实例上运行流行的 DNN 模型来广泛表征它们。具体来说,它估计了两种类型的通信停顿,即互连和网络停顿,它们在 DDL 执行时间中起主导作用。 Stash 是在之前的工作 DS-analyzer 之上实现的,DS-analyzer 仅计算 CPU 和磁盘停顿。通过详细的停顿特征,我们列出了公共云 GPU 实例的优点和缺点,以帮助用户做出明智的决策。我们的表征结果表明,更昂贵的 GPU 实例可能不是所有 DNN 模型的最佳性能,并且 AWS 有时可能无法最佳地分配硬件互连资源。具体来说,机器内互连可能会带来高达 90% 的 DNN 训练时间的通信开销,并且与单个实例上的训练相比,网络连接的实例可能会出现高达 5 倍的速度减慢。此外,(iii) 我们还模拟了 DNN 宏观特征的影响,例如层数和梯度数量对通信停顿的影响,最后,(iv) 我们简要讨论了与现有工作的成本比较。
Deep neural networks (DNNs) are increasingly popular owing to their ability to solve complex problems such as image recognition, autonomous driving, and natural language processing. Their growing complexity coupled with the use of larger volumes of training data (to achieve acceptable accuracy) has warranted the use of GPUs and other accelerators. Such accelerators are typically expensive, with users having to pay a high upfront cost to acquire them. For infrequent use, users can, instead, leverage the public cloud to mitigate the high acquisition cost. However, with the wide diversity of hardware instances (particularly GPU instances) available in public cloud, it becomes challenging for a user to make an appropriate choice from a cost/performance standpoint. In this work, we try to address this problem by (i) introducing a comprehensive distributed deep learning (DDL) profiler Stash, which determines the various execution stalls that DDL suffers from, and (ii) using Stash to extensively characterize various public cloud GPU instances by running popular DNN models on them. Specifically, it estimates two types of communication stalls, namely, interconnect and network stalls, that play a dominant role in DDL execution time. Stash is implemented on top of prior work, DS-analyzer, that computes only the CPU and disk stalls. Using our detailed stall characterization, we list the advantages and shortcomings of public cloud GPU instances for users to help them make an informed decision(s). Our characterization results indicate that the more expensive GPU instances may not be the most performant for all DNN models and that AWS can sometimes sub-optimally allocate hardware interconnect resources. Specifically, the intra-machine interconnect can introduce communication overheads of up to 90% of DNN training time and the network-connected instances can suffer from up to 5× slowdown compared to training on a single instance. Furthermore, (iii) we also model the impact of DNN macroscopic features such as the number of layers and the number of gradients on communication stalls, and finally, (iv) we briefly discuss a cost comparison with existing work.