Artemis: Automatic Runtime Tuning of Parallel Execution Parameters Using Machine Learning

Artemis: Automatic Runtime Tuning of Parallel Execution Parameters Using Machine Learning
复制标题

Artemis:使用机器学习自动运行时调整并行执行参数

DOI:
10.1007/978-3-030-78713-4_24
复制
发表时间:
2021
期刊:
The International Journal of High Performance Computing Applications
影响因子:
--
通讯作者:
T. Gamblin
T. Gamblin
中科院分区:
--
文献类型:
--
作者:
Chad Wood;G. Georgakoudis;D. Beckingsale;David Poliakoff;Alfredo Giménez;K. Huck;A. Malony;T. Gamblin

文献摘要

被引文献

相似文献

.可移植的并行编程模型提供了高性能和高生产力的潜力,但是它们带有大量的运行时参数,这些参数可能会对执行性能产生重大影响。选择这些参数的最佳集合是不平凡的,使得HPC应用在不同的系统环境和不同的输入数据集上表现良好,而不需要耗时的参数探索或主要的算法调整。我们提出了Artemis,一种使用机器学习进行在线,反馈驱动,自动参数调整的方法,该方法可推广,适合集成到高性能代码中。Artemis在运行时监视执行,并创建自适应模型来调整执行参数,同时在应用程序开发和运行时开销方面具有最低限度的侵入性。我们通过优化三个HPC代理应用程序(Cleverleaf、LULESH和Kokkos Kernels SpMV)的执行时间来展示Artemis的高效性。评估表明,Artemis选择最佳执行策略的准确率超过85%,具有不到9%的适度监控开销,并将执行速度提高了47%,尽管它的运行时开销。
. Portable parallel programming models provide the potential for high performance and productivity, however they come with a multitude of runtime parameters that can have significant impact on execution performance. Selecting the optimal set of those parameters is non-trivial, so that HPC applications perform well in different system environments and on different input data sets, without the need of time consuming parameter exploration or major algorithmic adjustments. We present Artemis, a method for online, feedback-driven, automatic parameter tuning using machine learning that is generalizable and suitable for integration into high-performance codes. Artemis monitors execution at runtime and creates adaptive models for tuning execution parameters, while being minimally invasive in application development and runtime overhead. We demonstrate the effectiveness of Artemis by optimizing the execution times of three HPC proxy applications: Cleverleaf, LULESH, and Kokkos Kernels SpMV. Evaluation shows that Artemis selects the optimal execution policy with over 85% accuracy, has modest monitoring overhead of less than 9%, and increases execution speed by up to 47% despite its runtime overhead.