OpenCL Performance on the Intel Heterogeneous Architecture Research Platform

OpenCL Performance on the Intel Heterogeneous Architecture Research Platform
复制标题

DOI:
10.1109/hpec43674.2020.9286213
复制
发表时间:
2020-09
期刊:
2020 IEEE High Performance Extreme Computing Conference (HPEC)
影响因子:
--
通讯作者:
Steven Harris;R. Chamberlain;Christopher D. Gill
Steven Harris;R. Chamberlain;Christopher D. Gill
中科院分区:
其他
文献类型:
--
作者:
Steven Harris;R. Chamberlain;Christopher D. Gill

文献摘要

相似文献

矩阵乘法的基本运算在众多学科中无处不在。然而,确定矩阵乘法的新优化仍然与新兴的硬件体系结构和不同的系统相关。OpenCL等框架支持在现有系统上进行计算协调,其使用英特尔高级合成编译器的可用性允许用户使用C/C++为可重新配置的硬件构建新的设计。使用HARPv2作为探索的载体,我们调查了几种传统的矩阵乘法优化的实用性,以更好地理解OpenCL的性能可移植性以及这种优化对缓存一致性异质体系结构的影响。我们的结果为最佳实践的适用性提供了有针对性的见解,这些最佳实践是为现有体系结构设计的,当用于新兴的异类系统时。
The fundamental operation of matrix multiplication is ubiquitous across a myriad of disciplines. Yet, the identification of new optimizations for matrix multiplication remains relevant for emerging hardware architectures and heterogeneous systems. Frameworks such as OpenCL enable computation orchestration on existing systems, and its availability using the Intel High Level Synthesis compiler allows users to architect new designs for reconfigurable hardware using C/C++. Using the HARPv2 as a vehicle for exploration, we investigate the utility of several traditional matrix multiplication optimizations to better understand the performance portability of OpenCL and the implications for such optimizations on cache coherent heterogeneous architectures. Our results give targeted insights into the applicability of best practices that were designed for existing architectures when used on emerging heterogeneous systems.