A 3 D-Stacked Memory Manycore Stencil Accelerator System

A 3 D-Stacked Memory Manycore Stencil Accelerator System
复制标题

3D堆叠内存众核模板加速器系统

DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
F. Franchetti
F. Franchetti
中科院分区:
--
文献类型:
--
作者:
Jiyuan Zhang;Tze Meng Low;Qi Guo;F. Franchetti

文献摘要

被引文献

相似文献

摘要空间运算是一类重要的科学计算内核,在科学模拟和图像处理中普遍存在。这类计算的关键特征是它们具有低操作强度,即,存储器访问的数量与其执行的浮点运算的数量的比率很高。因此,在通用计算系统上实现的模板操作的性能受到存储器带宽的限制。诸如3D堆叠存储器之类的技术可以提供比传统存储器系统大得多的带宽,并且可以增强诸如模板内核之类的存储器密集型计算的性能。在本文中,我们利用这种3D堆叠存储器技术来设计一个加速器的模板计算。我们表明,最好的效率,需要找到计算和内存访问之间的平衡,以保持所有组件始终忙碌。我们实现这一点,探讨如何阻塞和缓存方案来控制计算内存的比例。最后,我们确定最佳的设计点,最大限度地提高性能。
Stencil operations are an important class of scientific computational kernels that are pervasive in scientific simulations as well as in image processing. A key characteristic of this class of computation is that they have a low operational intensity, i.e., the ratio of the number of memory accesses to the number of floating point operations it performs is high. As a result, the performance of stencil operations implemented on general purpose computing systems is bounded by the memory bandwidth. Technologies such as 3D stacked memory can provide substantially more bandwidth than conventional memory systems and can enhance the performance of memory intensive computations like stencil kernels. In this paper, we leverage this 3D stacked memory technology to design an accelerator for stencil computations. We show that for the best efficiency one needs to find the balance between computation and memory accesses to keep all components consistently busy. We achieve this by exploring how blocking and caching schemes to control the compute-to-memory ratio. Finally, we identify optimal design points that maximize performance.