Memory access optimization in compilation for coarse-grained reconfigurable architectures

Memory access optimization in compilation for coarse-grained reconfigurable architectures
复制标题

粗粒度可重构架构编译中的内存访问优化

DOI:
10.1145/2003695.2003702
复制
发表时间:
2011
期刊:
ACM Trans. Design Autom. Electr. Syst.
影响因子:
--
通讯作者:
Y. Paek
Y. Paek
中科院分区:
--
文献类型:
--
作者:
Yongjoo Kim;Jongeun Lee;Aviral Shrivastava;Y. Paek

文献摘要

被引文献

相似文献

粗粒度可重构架构 (CGRA) 保证了高性能和高功效。他们通过保持硬件极其简单并将复杂性转移到应用程序映射来实现这一承诺。一项主要挑战来自数据映射的形式。出于功效和复杂性的考虑,CGRA 使用多组本地内存,并且一行 PE 共享内存访问。为了使每行PE能够访问任意存储体,在PE生成的存储请求和本地存储器的存储体之间存在硬件仲裁器。然而,一个基本的限制仍然存在,即一个存储体不能同时被两个不同的 PE 访问。我们建议通过将应用程序操作映射到 PE 并将数据映射到内存库来应对这一挑战,从而避免此类冲突。为了进一步提高多存储体内存的性能,我们提出了针对 CGRA 映射的编译器优化,以通过利用数据重用来减少内存操作的数量。我们对来自多媒体基准的内核的实验结果表明,与内存不感知的调度程序相比,我们的本地内存感知编译方法可以生成性能提高高达 53% 的映射(平均 26%)。
Coarse-grained reconfigurable architectures (CGRAs) promise high performance at high power efficiency. They fulfil this promise by keeping the hardware extremely simple, and moving the complexity to application mapping. One major challenge comes in the form of data mapping. For reasons of power-efficiency and complexity, CGRAs use multibank local memory, and a row of PEs share memory access. In order for each row of the PEs to access any memory bank, there is a hardware arbiter between the memory requests generated by the PEs and the banks of the local memory. However, a fundamental restriction remains in that a bank cannot be accessed by two different PEs at the same time. We propose to meet this challenge by mapping application operations onto PEs and data into memory banks in a way that avoids such conflicts. To further improve performance on multibank memories, we propose a compiler optimization for CGRA mapping to reduce the number of memory operations by exploiting data reuse. Our experimental results on kernels from multimedia benchmarks demonstrate that our local memory-aware compilation approach can generate mappings that are up to 53% better in performance (26% on average) compared to a memory-unaware scheduler.