Evaluating vector data type usage in OpenCL kernels

Evaluating vector data type usage in OpenCL kernels
复制标题

DOI:
10.1002/cpe.3424
复制
发表时间:
2015-12
期刊:
Concurrency and Computation: Practice and Experience
影响因子:
--
通讯作者:
Jianbin Fang;A. Varbanescu;Xiangke Liao;H. Sips
Jianbin Fang;A. Varbanescu;Xiangke Liao;H. Sips
中科院分区:
其他
文献类型:
--
作者:
Jianbin Fang;A. Varbanescu;Xiangke Liao;H. Sips

文献摘要

被引文献

相似文献

开放计算语言(OpenCL)是一种开放的、功能可移植的编程模型,适用于大范围的高度并行处理器。为了向用户提供对底层平台的访问,OpenCL明确支持本地内存和矢量数据类型(vdt)等特性。然而,这些通常是低级的、特定于硬件的功能,这可能会损害不同平台上的性能。本文以vdt为研究对象,对其应用进行了系统的研究。首先,我们提出了在OpenCL内核中使用vdt的两种不同方法(vdt间和vdt内),并展示了如何将标量OpenCL内核转换为矢量化内核。在获得矢量化代码后,我们用两种类型的基准评估了使用vdt的性能影响:微观基准和宏观基准。通过微基准测试,我们研究了vdt的执行模型和编译器辅助矢量器在五种设备上的作用。通过宏观基准测试,我们探索了使用vdt之前和之后内存访问模式的变化,以及由此产生的性能影响。我们的评估不仅提供了对OpenCL的vdt如何映射到不同处理器上的见解,而且还表明使用这些数据类型会在计算和内存访问中引入变化。根据吸取的经验教训,我们将讨论如何在存在vdt的情况下处理性能可移植性。版权所有©2014 John Wiley & Sons, Ltd。
Open Computing Language (OpenCL) is an open, functionally portable programming model for a large range of highly parallel processors. To provide users with access to the underlying platforms, OpenCL has explicit support for features such as local memory and vector data types (VDTs). However, these are often low‐level, hardware‐specific features, which can be detrimental to performance on different platforms. In this paper, we focus on VDTs and investigate their usage in a systematic way. First, we propose two different approaches (inter‐vdt and intra‐vdt) to use VDTs in OpenCL kernels, and show how to translate scalar OpenCL kernels to vectorized ones. After obtaining vectorized code, we evaluate the performance effects of using VDTs with two types of benchmarks: micro‐benchmarks and macro‐benchmarks. With micro‐benchmarks, we study the execution model of VDTs and the role of the compiler‐aided vectorizer on five devices. With macro‐benchmarks, we explore the changes of memory access patterns before and after using VDTs, and the resulting performance impact. Not only our evaluation provides insights into how OpenCL's VDTs are mapped on different processors, but it also indicates that using such data types introduces changes in both computation and memory accesses. Based on the lessons learned, we discuss how to deal with performance portability in the presence of VDTs. Copyright © 2014 John Wiley & Sons, Ltd.