A comparison of GPU execution time prediction using machine learning and analytical modeling

A comparison of GPU execution time prediction using machine learning and analytical modeling
复制标题

DOI:
10.1109/nca.2016.7778637
复制
发表时间:
2016-10
期刊:
2016 IEEE 15th International Symposium on Network Computing and Applications (NCA)
影响因子:
--
通讯作者:
Marcos Amarís;R. Camargo;M. Dyab;A. Goldman;D. Trystram
Marcos Amarís;R. Camargo;M. Dyab;A. Goldman;D. Trystram
中科院分区:
其他
文献类型:
--
作者:
Marcos Amarís;R. Camargo;M. Dyab;A. Goldman;D. Trystram

文献摘要

被引文献

相似文献

如今,大多数高性能计算(HPC)平台具有不同的硬件资源(CPU、GPU、存储等)。图形处理单元(GPU)是专门用于加速向量运算的并行计算协处理器。对这些设备上的应用程序执行时间的预测是一个巨大的挑战,并且对于高效的作业调度是必不可少的。有不同的方法可以做到这一点,例如分析建模和机器学习技术。分析预测模型很有用,但需要手动包含体系结构和软件之间的交互,并且可能无法捕获GPU体系结构中的复杂交互。机器学习技术可以在不需要人工干预的情况下学习捕获这些交互,但可能需要大量的训练集。在本文中,我们比较了三种不同的机器学习方法:线性回归、支持向量机和随机森林,并使用基于BSP的分析模型来预测GPU应用程序的执行时间。作为机器学习算法的输入,我们使用了在9个不同的GPU上执行的9个不同应用程序的分析信息。我们表明,机器学习方法为不同的情况提供了合理的预测。尽管预测不如分析模型,但它们不需要应用程序代码、硬件特征或显式建模的详细知识。因此,每当具有简档信息的数据库可用或可以生成时,机器学习技术对于在包含GPU的异类架构上部署调度应用的自动化在线性能预测是有用的。
Today, most high-performance computing (HPC) platforms have heterogeneous hardware resources (CPUs, GPUs, storage, etc.) A Graphics Processing Unit (GPU) is a parallel computing coprocessor specialized in accelerating vector operations. The prediction of application execution times over these devices is a great challenge and is essential for efficient job scheduling. There are different approaches to do this, such as analytical modeling and machine learning techniques. Analytic predictive models are useful, but require manual inclusion of interactions between architecture and software, and may not capture the complex interactions in GPU architectures. Machine learning techniques can learn to capture these interactions without manual intervention, but may require large training sets. In this paper, we compare three different machine learning approaches: linear regression, support vector machines and random forests with a BSP-based analytical model, to predict the execution time of GPU applications. As input to the machine learning algorithms, we use profiling information from 9 different applications executed over 9 different GPUs. We show that machine learning approaches provide reasonable predictions for different cases. Although the predictions were inferior to the analytical model, they required no detailed knowledge of application code, hardware characteristics or explicit modeling. Consequently, whenever a database with profile information is available or can be generated, machine learning techniques can be useful for deploying automated on-line performance prediction for scheduling applications on heterogeneous architectures containing GPUs.