Automatic On-chip Memory Minimization for Data Reuse

Automatic On-chip Memory Minimization for Data Reuse
复制标题

自动片上内存最小化以实现数据重用

DOI:
--
复制
发表时间:
2007
期刊:
IEEE Symposium on Field-Programmable Custom Computing Machines
影响因子:
--
通讯作者:
P. Cheung
P. Cheung
中科院分区:
--
文献类型:
--
作者:
Qiang Liu;G. Constantinides;K. Masselos;P. Cheung

文献摘要

被引文献

相似文献

基于现场可编程门阵列的计算引擎由于具有高度的灵活性和并行性,已经成为实现计算密集型应用的一种很有前途的选择。然而,当试图在FPGA上加速应用程序时,需要克服的主要障碍之一是片外通信的瓶颈,通常是与大容量存储器的通信。通常,在编译时知道相同的数据项被多次访问,因此可以从大的片外RAM一次加载到稀缺的片上RAM上,从而缓解了这一瓶颈。本文讨论如何自动导出地址映射,以减少给定存储器访问模式所需的片上存储器的大小。实验结果表明,在实践中,我们的方法将片上存储需求降到了最低,与幼稚的方法相比,某些基准测试的片上存储大小减少了40倍(平均为10倍)。同时,与这种方法相比,对于这些基准,没有观察到时钟周期损失或控制逻辑面积的增加。
FPGA-based computing engines have become a promising option for the implementation of computationally intensive applications due to high flexibility and parallelism. However, one of the main obstacles to overcome when trying to accelerate an application on an FPGA is the bottleneck in off-chip communication, typically to large memories. Often it is known at compile-time that the same data item is accessed many times, and as a result can be loaded once from large off-chip RAM onto scarce on-chip RAM, alleviating this bottleneck. This paper addresses how to automatically derive an address mapping that reduces the size of the required on-chip memory for a given memory access pattern. Experimental results demonstrate that, in practice, our approach reduces on-chip storage requirements to the minimum, corresponding to a reduction in on-chip memory size of up to 40times (average 10times) for some benchmarks compared to a naive approach. At the same time, no clock period penalty or increase in control logic area compared to this approach is observed for these benchmarks.