Benchmarking the Performance of Accelerators on National Cyberinfrastructure Resources for Artificial Intelligence / Machine Learning Workloads

Benchmarking the Performance of Accelerators on National Cyberinfrastructure Resources for Artificial Intelligence / Machine Learning Workloads
复制标题

DOI:
10.1145/3491418.3530772
复制
发表时间:
2022-07
期刊:
Practice and Experience in Advanced Research Computing
影响因子:
--
通讯作者:
Abhinand Nasari;Hieu Hanh Le;Richard Lawrence;Zhenhua He;Xin Yang;Mario Krell;A. Tsyplikhin;M. Tatineni;Tim Cockerill;Lisa M. Perez;Dhruva K. Chakravorty;Honggao Liu
Abhinand Nasari;Hieu Hanh Le;Richard Lawrence;Zhenhua He;Xin Yang;Mario Krell;A. Tsyplikhin;M. Tatineni;Tim Cockerill;Lisa M. Perez;Dhruva K. Chakravorty;Honggao Liu
中科院分区:
其他
文献类型:
--
作者:
Abhinand Nasari;Hieu Hanh Le;Richard Lawrence;Zhenhua He;Xin Yang;Mario Krell;A. Tsyplikhin;M. Tatineni;Tim Cockerill;Lisa M. Perez;Dhruva K. Chakravorty;Honggao Liu

文献摘要

相似文献

即将到来的区域和国家科学基金会(NSF)资助的网络基础设施(CI)资源将为研究人员提供在加速器上运行人工智能/机器学习(AI/ML)工作流程的机会。为了有效地利用这一新兴的CI丰富的景观,研究人员需要广泛的基准数据,以最大限度地提高性能,并将其工作流程映射到适当的架构。这些数据将进一步帮助CI管理员,NSF项目官员和CI分配评审员对CI资源分配做出明智的决定。在这里,我们通过运行常见AI/ML模型的训练基准来比较两种非常不同的架构的性能:常用的图形处理单元(GPU)和新一代智能处理单元(IPU)。我们利用软件堆栈的成熟度以及这些平台之间的易迁移性来了解两种架构的性能和扩展性是相似的。然而,探索训练参数(如批量大小)发现,由于内存处理结构,IPU可以在较小的批量大小下高效运行,而GPU则可以从大批量中受益,从而在神经网络训练和推理中提取足够的并行性。正如本文所讨论的,这带来了不同的优点和缺点。因此,推理延迟,固有并行性和模型准确性的考虑将在研究人员选择这些架构时发挥作用。这些选择的影响,一个代表性的图像压缩模型系统进行了讨论。
Upcoming regional and National Science Foundation (NSF)-funded Cyberinfrastructure (CI) resources will give researchers opportunities to run their artificial intelligence / machine learning (AI/ML) workflows on accelerators. To effectively leverage this burgeoning CI-rich landscape, researchers need extensive benchmark data to maximize performance gains and map their workflows to appropriate architectures. This data will further assist CI administrators, NSF program officers, and CI allocation-reviewers make informed determinations on CI-resource allocations. Here, we compare the performance of two very different architectures: the commonly used Graphical Processing Units (GPUs) and the new generation of Intelligence Processing Units (IPUs), by running training benchmarks of common AI/ML models. We leverage the maturity of software stacks, and the ease of migration among these platforms to learn that performance and scaling are similar for both architectures. Exploring training parameters, such as batch size, however finds that owing to memory processing structures, IPUs run efficiently with smaller batch sizes, while GPUs benefit from large batch sizes to extract sufficient parallelism in neural network training and inference. This comes with different advantages and disadvantages as discussed in this paper.As such considerations of inference latency, inherent parallelism and model accuracy will play a role in researcher selection of these architectures. The impact of these choices on a representative image compression model system is discussed.