MLPerf™ HPC: A Holistic Benchmark Suite for Scientific Machine Learning on HPC Systems

MLPerf™ HPC: A Holistic Benchmark Suite for Scientific Machine Learning on HPC Systems
复制标题

MLPerf™ HPC:HPC 系统上科学机器学习的整体基准套件

DOI:
10.1109/mlhpc54614.2021.00009
复制
发表时间:
2021
期刊:
2021 IEEE/ACM Workshop on Machine Learning in High Performance Computing Environments (MLHPC
影响因子:
--
通讯作者:
Mattson, Peter
Mattson, Peter
中科院分区:
--
文献类型:
--
作者:
Farrell, Steven;Emani, Murali;Balma, Jacob;Drescher, Lukas;Drozd, Aleksandr;Fink, Andreas;Fox, Geoffrey;Kanter, David;Kurth, Thorsten;Mattson, Peter

文献摘要

相似文献

科学界越来越多地在其应用程序中采用机器学习和深度学习模型,以加速科学见解。高性能计算系统正在以丰富多样的硬件资源和大规模横向扩展能力推动性能的前沿。我们迫切需要了解代表真实世界科学用例的机器学习应用程序的公平有效的基准测试。MLPerf™是一个社区驱动的标准,用于对机器学习工作负载进行基准测试,专注于端到端性能指标。在本文中,我们介绍了MLPerf HPC,这是由MLCommons™协会驱动的大型科学机器学习训练应用程序的基准套件。我们展示了第一轮提交的结果,其中包括一些世界上最大的HPC系统。我们开发了一个系统的框架,他们的联合分析,并比较他们的数据分期,算法收敛和计算性能。因此,我们获得了对不同子系统的优化的定量理解,例如数据的分级和节点加载,计算单元利用率和通信调度,通过系统扩展实现整体(端到端)性能的提高。值得注意的是,我们的分析显示了数据集大小,系统的内存层次结构和训练收敛之间的规模依赖性相互作用,强调了近计算存储的重要性。为了克服大批量数据并行可扩展性的挑战,我们讨论了特定的学习技术和混合数据和模型的并行性,是有效的大型系统。最后,我们通过表征每个基准测试的低级别内存,I/O和网络行为参数扩展车顶性能模型在未来的几轮。
Scientific communities are increasingly adopting machine learning and deep learning models in their applications to accelerate scientific insights. High performance computing systems are pushing the frontiers of performance with a rich diversity of hardware resources and massive scale-out capabilities. There is a critical need to understand fair and effective benchmarking of machine learning applications that are representative of real-world scientific use cases. MLPerf™is a community-driven standard to benchmark machine learning workloads, focusing on end-to-end performance metrics. In this paper, we introduce MLPerf HPC, a benchmark suite of large-scale scientific machine learning training applications, driven by the MLCommons™Association. We present the results from the first submission round including a diverse set of some of the world’s largest HPC systems. We develop a systematic framework for their joint analysis and compare them in terms of data staging, algorithmic convergence and compute performance. As a result, we gain a quantitative understanding of optimizations on different subsystems such as staging and on-node loading of data, compute-unit utilization and communication scheduling enabling overall(end-to-end) performance improvements through system scaling. Notably, our analysis shows a scale-dependent interplay between the dataset size, a system’s memory hierarchy and training convergence that underlines the importance of near-compute storage. To overcome the data-parallel scalability challenge at large batch-sizes, we discuss specific learning techniques and hybrid data-and-model parallelism that are effective on large systems. We conclude by characterizing each benchmark with respect to low-level memory, I/O and network behaviour to parameterize extended roofline performance models in future rounds.