Performance of Distributed Deep Learning Workloads on a Composable Cyberinfrastructure

Performance of Distributed Deep Learning Workloads on a Composable Cyberinfrastructure
复制标题

可组合网络基础设施上分布式深度学习工作负载的性能

DOI:
10.1145/3569951.3593601
复制
发表时间:
2023
期刊:
USA.
影响因子:
--
通讯作者:
Liu, Honggao
Liu, Honggao
中科院分区:
--
文献类型:
--
作者:
He, Zhenhua;Saluja, Aditi;Lawrence, Richard;Chakravorty, Dhruva;Dang, Francis;Perez, Lisa;Liu, Honggao

文献摘要

参考文献

被引文献

相似文献

下一代计算系统可能依赖可动态重新配置和定制的分类资源,以支持需要不同网络基础设施(CI)技术的科学和工程工作流程。这些资源将包括内存、加速器、协处理器等技术。这将代表高性能计算(HPC)的重大转变,不同于目前将这些资源永久连接到一台服务器的典型集群模型。虽然用分解的资源组合硬件框架是有希望的,但我们需要了解如何将工作流放置在这些资源上,并评估这种方法对工作流性能的影响。为了开发这个知识框架,我们研究了深度学习工作负载在支持GPU的可组合计算平台和传统HPC计算平台上的适用性和性能。这里给出了在这些HPC环境中使用Horovod框架和TensorFlow和PyTorch模型执行的测试结果。
The next generation of computing systems are likely to rely on disaggregated resources that can be dynamically reconfigured and customized for researchers to support scientific and engineering workflows that require different cyberinfrastructure (CI) technologies. These resources would include memory, accelerators, co-processors among other technologies. This would represent a significant shift in High Performance Computing (HPC) from the now typical model of clusters that have these resources permanently connected to a single server. While composing hardware frameworks with disaggregated resources holds promise, we need to understand how to situate workflows on these resources and evaluate the impact of this approach on workflow performance against “traditional” clusters.  Toward developing this knowledge framework, we study the applicability and performance of deep learning workloads on GPU-enabled composable and traditional HPC computing platforms. Results from tests performed using the Horovod framework with TensorFlow and PyTorch models on these HPC environments are presented here.
第一届可组合系统研讨会 (COMPSYS 2022)
DOI: --
发表时间: 2022
期刊: IEEE International Symposium on Parallel & Distributed Processing, Workshops and Phd Forum
影响因子: --
作者:
A. Nichols
通讯作者: A. Nichols