Demystifying the MLPerf Training Benchmark Suite

Demystifying the MLPerf Training Benchmark Suite
复制标题

DOI:
10.1109/ispass48437.2020.00013
复制
发表时间:
2020-08
期刊:
2020 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS)
影响因子:
--
通讯作者:
Snehil Verma;Qinzhe Wu;Bagus Hanindhito;Gunjan Jha;E. John;R. Radhakrishnan;L. John
Snehil Verma;Qinzhe Wu;Bagus Hanindhito;Gunjan Jha;E. John;R. Radhakrishnan;L. John
中科院分区:
其他
文献类型:
--
作者:
Snehil Verma;Qinzhe Wu;Bagus Hanindhito;Gunjan Jha;E. John;R. Radhakrishnan;L. John

文献摘要

相似文献

MLPerf是一个新兴的机器学习基准套件,致力于覆盖广泛的机器学习应用。我们对MLPerf基准测试的特点以及它们与以前的深度学习基准测试(如DAWNBench和DeepBench)的区别进行了研究。MLPerf基准测试显示出中等高的每秒内存事务和中等高的计算速率,而DAWNBench创建了具有低内存事务速率的高计算基准测试,DeepBench提供了低计算速率基准测试。我们还观察到,各种MLPerf基准测试具有独特的功能,可以揭示系统中的各种瓶颈。我们还观察到MLPerf模型中缩放效率的变化。不同模型表现出的差异突出了智能调度策略对多GPU训练的重要性。另一个观察结果是,多GPU系统中GPU之间的专用低延迟互连对于优化分布式深度学习训练至关重要。此外,主机CPU利用率随着用于训练的GPU数量的增加而增加。为了证实之前的工作,我们还观察并量化了使用Tensor Cores进行混合精度训练可能带来的改进。
MLPerf, an emerging machine learning benchmark suite, strives to cover a broad range of machine learning applications. We present a study on the characteristics of MLPerf benchmarks and how they differ from previous deep learning benchmarks such as DAWNBench and DeepBench. MLPerf benchmarks are seen to exhibit moderately high memory transactions per second and moderately high compute rates, while DAWNBench creates a high-compute benchmark with low memory transaction rate, and DeepBench provides low compute rate benchmarks. We also observe that the various MLPerf benchmarks possess unique features that allow unveiling various bottlenecks in systems. We also observe variation in scaling efficiency across the MLPerf models. The variation exhibited by the different models highlight the importance of smart scheduling strategies for multi-GPU training. Another observation is that dedicated low latency interconnect between GPUs in multi-GPU systems is crucial for optimal distributed deep learning training. Furthermore, host CPU utilization increases with an increase in the number of GPUs used for training. Corroborating prior work, we also observe and quantify improvements possible by mixed-precision training using Tensor Cores.