A High Throughput FPGA-based Floating Point Conjugate Gradient Implementation

A High Throughput FPGA-based Floating Point Conjugate Gradient Implementation
复制标题

基于 FPGA 的高吞吐量浮点共轭梯度实现

DOI:
10.1007/978-3-540-78610-8_10
复制
发表时间:
2008
期刊:
The American Mathematical Monthly
影响因子:
--
通讯作者:
G. Constantinides
G. Constantinides
中科院分区:
--
文献类型:
--
作者:
Antonio Roldao Lopes;G. Constantinides

文献摘要

被引文献

相似文献

随着现场可编程门阵列(fpga)达到超过数百万等效门的容量,加速浮点科学计算应用成为可能。在科学计算中常见的一类计算是线性方程组的解。一种在软件中被证明是非常有效和健壮的求解方法是共轭梯度算法。本文提出了一种并行的硬件共轭梯度实现。该实现特别适合于加速多个中小型密集线性方程组。通过并行化,可以将每次迭代的计算时间从i¾?(n2)个周期用于软件实现到i¾?(n)。I/O需求是可伸缩的,并且随着矩阵阶数的增加收敛到一个恒定值。在VirtexII-6000上的结果显示持续性能为5 GFLOPS,而在Virtex5-330上的预测结果显示持续性能为35 GFLOPS。前者的结果与高端cpu相当,而后者则代表了显著的加速。
As Field Programmable Gate Arrays (FPGAs) have reached capacities beyond millions of equivalent gates, it becomes possible to accelerate floating-point scientific computing applications. One type of calculation that is commonplace in scientific computation is the solution of systems of linear equations. A method that has proven in software to be very efficient and robust for finding such solutions is the Conjugate Gradient algorithm. In this paper we present a parallel hardware Conjugate Gradient implementation. The implementation is particularly suited for accelerating multiple small to medium sized dense systems of linear equations. Through parallelization it is possible to convert the computation time per iteration for an order nmatrix from i¾?(n2) cycles for a software implementation to i¾?(n). I/O requirements are scalable and converge to a constant value with the increase of matrix order. Results on a VirtexII-6000 demonstrate sustained performance of 5 GFLOPS and projected results on a Virtex5-330 indicate sustained performance of 35 GFLOPS. The former result is comparable to high-end CPUs, whereas the latter represents a significant speedup.