Portable and scalable FPGA-based acceleration of a direct linear system solver

Portable and scalable FPGA-based acceleration of a direct linear system solver
复制标题

DOI:
10.1145/2133352.2133358
复制
发表时间:
2008-12
期刊:
2008 International Conference on Field-Programmable Technology
影响因子:
--
通讯作者:
Wei Zhang-;Vaughn Betz;Jonathan Rose
Wei Zhang-;Vaughn Betz;Jonathan Rose
中科院分区:
其他
文献类型:
--
作者:
Wei Zhang-;Vaughn Betz;Jonathan Rose

文献摘要

被引文献

相似文献

FPGA正在成为一个有吸引力的平台,用于加速许多计算,包括科学应用。然而,它们的采用受到FPGA设计的大开发成本和短寿命的限制。我们相信,如果有硬件库可以移植到任何FPGA上,并且性能可以随FPGA的资源扩展,那么基于FPGA的科学计算将变得更加实用。为了说明这个想法,我们实现了一个常见的超级计算库函数:用于求解线性系统的LU分解方法。本文讨论的问题,使设计的可移植性和可扩展性。设计是自动生成的,以匹配FPGA Apsilas的能力和外部存储器通过使用参数。我们将FPGA上的设计性能与单个处理器内核进行了比较,发现其执行速度快2.2倍,每次计算消耗的能量少5倍。
FPGAs are becoming an attractive platform for accelerating many computations including scientific applications. However, their adoption has been limited by the large development cost and short life span of FPGA designs. We believe that FPGA-based scientific computation would become far more practical if there were hardware libraries that were portable to any FPGA with performance that could scale with the resources of the FPGA. To illustrate this idea we have implemented one common supercomputing library function: the LU factorization method for solving linear systems. This paper discusses issues in making the design both portable and scalable. The design is automatically generated to match the FPGApsilas capabilities and external memory through the use of parameters. We compared the performance of the design on the FPGA to a single processor core and found that it performs 2.2 times faster, and that the energy dissipated per computation is a factor 5 times less.