HeteroMap: A Runtime Performance Predictor for Efficient Processing of Graph Analytics on Heterogeneous Multi-Accelerators

HeteroMap: A Runtime Performance Predictor for Efficient Processing of Graph Analytics on Heterogeneous Multi-Accelerators
复制标题

DOI:
10.1109/ispass.2019.00039
复制
发表时间:
2019-03
期刊:
2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS)
影响因子:
--
通讯作者:
Masab Ahmad;H. Dogan;Christopher J. Michael;O. Khan
Masab Ahmad;H. Dogan;Christopher J. Michael;O. Khan
中科院分区:
其他
文献类型:
--
作者:
Masab Ahmad;H. Dogan;Christopher J. Michael;O. Khan

文献摘要

被引文献

相似文献

随着数量不断增加的数据和输入变化,便携式性能变得越来越难以利用当今的体系结构。计算设置利用单芯片处理器,例如GPU或大规模多头器进行图形分析。当利用GPU的较高并发性和带宽时,某些算法输入组合的性能更有效,而另一些算法组合则具有更强的数据缓存功能。在选定的加速器中,架构选择还会发生,其中需要确定诸如线程和线程放置之类的变量以达到最佳性能。本文提出了针对异质并行体系结构的性能预测范式,其中将多个不同的加速器集成到操作高性能计算设置中。该预测变量旨在通过使用图基准和输入特性来利用异质集成加速器内部和跨越异质的集成加速器的潜在并发变化来提高图形处理效率。评估表明,智能和实时选择近乎最佳的并发选择可提供5%至3.8 x的性能优势,而能量益处平均比传统的单加速器设置约为2.4 x。
With the ever-increasing amount of data and input variations, portable performance is becoming harder to exploit on today's architectures. Computational setups utilize single-chip processors, such as GPUs or large-scale multicores for graph analytics. Some algorithm-input combinations perform more efficiently when utilizing a GPU's higher concurrency and bandwidth, while others perform better with a multicore's stronger data caching capabilities. Architectural choices also occur within selected accelerators, where variables such as threading and thread placement need to be decided for optimal performance. This paper proposes a performance predictor paradigm for a heterogeneous parallel architecture where multiple disparate accelerators are integrated in an operational high performance computing setup. The predictor aims to improve graph processing efficiency by exploiting the underlying concurrency variations within and across the heterogeneous integrated accelerators using graph benchmark and input characteristics. The evaluation shows that intelligent and real-time selection of near-optimal concurrency choices provides performance benefits ranging from 5 % to 3.8 x, and an energy benefit averaging around 2.4 x over the traditional single-accelerator setup.