A Traversal Cache Framework for FPGA Acceleration of Pointer Data Structures: A Case Study on Barnes-Hut N-body Simulation

A Traversal Cache Framework for FPGA Acceleration of Pointer Data Structures: A Case Study on Barnes-Hut N-body Simulation
复制标题

一种用于 FPGA 指针数据结构加速的遍历缓存框架:Barnes-Hut N 体仿真案例研究

DOI:
--
复制
发表时间:
2009
期刊:
International Conference on Reconfigurable Computing and FPGAs
影响因子:
--
通讯作者:
G. Stitt
G. Stitt
中科院分区:
--
文献类型:
--
作者:
J. Coole;J. Wernsing;G. Stitt

文献摘要

被引文献

相似文献

大量研究表明,与微处理器相比,现场可编程门阵列(FPGA)通常可以实现较大的加速比。然而,FPGA的一个重要限制是对常规存储器访问模式的要求,这阻碍了它们在重要应用中的使用。Traffic缓存以前被引入,以提高不规则的内存访问模式,特别是那些遍历基于指针的数据结构的算法的FPGA实现的性能。然而,以前的遍历缓存的一个重要限制是,加速仅限于随着时间的推移频繁重复的遍历,从而阻止了没有重复的算法的加速,即使遍历之间的相似性很大。本文提出了一种新的框架,扩展遍历缓存,使性能提高,在这种情况下,并提供了额外的改进,通过减少内存访问和并行处理的多个遍历。最重要的是,我们表明,具有高度相似的遍历算法,遍历缓存框架实现了近似线性的内核加速与额外的面积,从而消除了通常与FPGA相关的内存带宽瓶颈。我们使用Barnes-Hut n体仿真案例研究评估该框架,显示Virtex 4 LX 100上的应用程序加速比从12倍到13.5倍不等,预计当今最大的FPGA上的加速比高达40倍。
Numerous studies have shown that field-programmable gate arrays (FPGAs) often achieve large speedups compared to microprocessors. However, one significant limitation of FPGAs that has prevented their use on important applications is the requirement for regular memory access patterns. Traversal caches were previously introduced to improve the performance of FPGA implementations of algorithms with irregular memory access patterns, especially those traversing pointer-based data structures. However, a significant limitation of previous traversal caches is that speedup was limited to traversals repeated frequently over time, thus preventing speedup for algorithms without repetition, even if the similarity between traversals was large. This paper presents a new framework that extends traversal caches to enable performance improvements in such cases and provides additional improvements through reduced memory accesses and parallel processing of multiple traversals. Most importantly, we show that, for algorithms with highly similar traversals, the traversal cache framework achieves approximately linear kernel speedup with additional area, thus eliminating the memory bandwidth bottleneck commonly associated with FPGAs. We evaluate the framework using a Barnes-Hut n-body simulation case study, showing application speedups ranging from 12x to 13.5x on a Virtex4 LX100 with projected speedups as high as 40x on today’s largest FPGAs.