A High Throughput FPGA-Based Floating Point Conjugate Gradient Implementation for Dense Matrices

A High Throughput FPGA-Based Floating Point Conjugate Gradient Implementation for Dense Matrices
复制标题

基于 FPGA 的高吞吐量密集矩阵浮点共轭梯度实现

DOI:
10.1145/1661438.1661439
复制
发表时间:
2010
影响因子:
2.3
通讯作者:
Roldao A
Roldao A
中科院分区:
计算机科学3区
文献类型:
--
作者:
Roldao A

文献摘要

参考文献

被引文献

相似文献

现代现场可编程门阵列(FPGA)的能力的最新发展显著地扩展了它们的应用。一个这样的领域是科学计算的加速,并且在科学计算中常见的一种类型的计算是线性方程组的解。在软件中已经证明对于找到这样的解决方案非常有效和鲁棒的方法是共轭梯度(CG)算法。在这篇文章中,我们提出了一个广泛的并行和深度流水线的硬件CG实现,针对现代FPGA架构。此实现特别适合于加速多个中小型密集线性方程组,可用作独立求解器或构建块来求解高阶系统。在这篇文章中,它表明,通过并行化,它是可能的转换每次迭代的计算时间为ordernmatrix从Θ(n2)时钟周期的微处理器上的FPGA上的Θ(n)。通过深度流水线,还可以并行解决多个问题,并最大限度地提高性能和效率。I/O的要求是可扩展的,收敛到一个恒定的值与矩阵阶数的增加。在现成的VirtexII-6000上的布局和布线后结果表明持续性能为5 GFlops,在Virtex 5 -330上的结果表明持续性能为35 GFlops。与高端CPU上运行的优化软件实现的比较表明,这种FPGA实现代表了至少一个数量级的显着加速。
Recent developments in the capacity of modern Field Programmable Gate Arrays (FPGAs) have significantly expanded their applications. One such field is the acceleration of scientific computation and one type of calculation that is commonplace in scientific computation is the solution of systems of linear equations. A method that has proven in software to be very efficient and robust for finding such solutions is the Conjugate Gradient (CG) algorithm. In this article we present a widely parallel and deeply pipelined hardware CG implementation, targeted at modern FPGA architectures. This implementation is particularly suited for accelerating multiple small-to-medium-sized dense systems of linear equations and can be used as a stand-alone solver or as building block to solve higher-order systems. In this article it is shown that through parallelization it is possible to convert the computation time per iteration for an ordernmatrix fromΘ(n2) clock cycles on a microprocessor toΘ(n) on a FPGA. Through deep pipelining it is also possible to solve several problems in parallel and maximize both performance and efficiency. I/O requirements are shown to be scalable and convergent to a constant value with the increase of matrix order. Post place-and-route results on a readily available VirtexII-6000 demonstrate sustained performance of 5 GFlops, and results on a Virtex5-330 indicate sustained performance of 35 GFlops. A comparison with an optimized software implementation running on a high-end CPU demonstrate that this FPGA implementation represents a significant speedup of at least an order of magnitude.
DOI: 10.1007/978-3-540-78610-8_10
发表时间: 2008
期刊: The American Mathematical Monthly
影响因子: --
作者:
Antonio Roldao Lopes;G. Constantinides
通讯作者: G. Constantinides
使用 Cholesky 分解在 CELL 处理器上求解线性方程组
DOI: 10.1109/tpds.2007.70813
发表时间: 2008
影响因子: 5.3
作者:
J. Kurzak;A. Buttari;J. Dongarra
通讯作者: J. Dongarra
FPGA 的浮点数据路径综合
DOI: --
发表时间: 2008
期刊: International Conference on Field-Programmable Logic and Applications
影响因子: --
作者:
M. Langhammer
通讯作者: M. Langhammer
DOI: --
发表时间: 2005
期刊:
影响因子: --
作者:
I. Pournara;C. Bouganis;G. Constantinides
通讯作者: G. Constantinides
DOI: --
发表时间: 2005
期刊: Parallel Processing and Applied Mathematics
影响因子: --
作者:
O. Maslennikov;V. Lepekha;A. Sergyienko
通讯作者: A. Sergyienko