Exploiting Memory Access Patterns to Improve Memory Performance in Data-Parallel Architectures

Exploiting Memory Access Patterns to Improve Memory Performance in Data-Parallel Architectures
复制标题

DOI:
10.1109/tpds.2010.107
复制
发表时间:
2011-01-01
影响因子:
5.3
通讯作者:
Kaeli, David
Kaeli, David
中科院分区:
计算机科学2区
文献类型:
--
作者:
Jang, Byunghyun;Schaa, Dana;Kaeli, David

文献摘要

被引文献

相似文献

GPU上的通用计算(GPGPU)的引入改变了并行计算的未来。这一现象的核心是大规模多线程、数据并行架构,这些架构拥有令人印象深刻的加速等级,提供低成本的超级计算以及有吸引力的功耗预算。即使考虑到GPGPU提供的众多好处,仍然存在一些障碍,延迟了这些架构的广泛采用。一个主要问题是数据并行架构中常见的内存子系统的异构性和分布式特性。应用程序加速高度依赖于能够有效地利用内存子系统,以便所有执行单元保持忙碌状态。在本文中,我们提出了提高数据并行体系结构上的应用程序的内存效率的技术,基于循环体中内存访问模式的分析和表征;我们通过数据转换来实现向量化,以使基于向量的体系结构(例如,例如,在一个实施例中,AMD GPU)和基于标量架构的算法内存选择(例如,例如,在一个实施例中,NVIDIA GPU)。我们证明了我们提出的方法的有效性,从广泛的基准套件的内核。对于所研究的基准内核,我们通过应用我们提出的方法实现了一致且显着的性能改进(在每个平台上分别比基线GPU实现高出11.4倍和13.5倍)。
The introduction of General-Purpose computation on GPUs (GPGPUs) has changed the landscape for the future of parallel computing. At the core of this phenomenon are massively multithreaded, data-parallel architectures possessing impressive acceleration ratings, offering low-cost supercomputing together with attractive power budgets. Even given the numerous benefits provided by GPGPUs, there remain a number of barriers that delay wider adoption of these architectures. One major issue is the heterogeneous and distributed nature of the memory subsystem commonly found on data-parallel architectures. Application acceleration is highly dependent on being able to utilize the memory subsystem effectively so that all execution units remain busy. In this paper, we present techniques for enhancing the memory efficiency of applications on data-parallel architectures, based on the analysis and characterization of memory access patterns in loop bodies; we target vectorization via data transformation to benefit vector-based architectures (e. g., AMD GPUs) and algorithmic memory selection for scalar-based architectures (e. g., NVIDIA GPUs). We demonstrate the effectiveness of our proposed methods with kernels from a wide range of benchmark suites. For the benchmark kernels studied, we achieve consistent and significant performance improvements (up to 11.4 x and 13.5 x over baseline GPU implementations on each platform, respectively) by applying our proposed methodology.