A High Throughput FPGA-Based Floating Point Conjugate Gradient Implementation for Dense Matrices
A High Throughput FPGA-Based Floating Point Conjugate Gradient Implementation for Dense Matrices
复制标题
基于 FPGA 的高吞吐量密集矩阵浮点共轭梯度实现
DOI:
10.1145/1661438.1661439
复制
发表时间:
2010
影响因子:
2.3
通讯作者:
Roldao A
中科院分区:
文献类型:
--
作者:
Roldao A
Recent developments in the capacity of modern Field Programmable Gate Arrays (FPGAs) have significantly expanded their applications. One such field is the acceleration of scientific computation and one type of calculation that is commonplace in scientific computation is the solution of systems of linear equations. A method that has proven in software to be very efficient and robust for finding such solutions is the Conjugate Gradient (CG) algorithm. In this article we present a widely parallel and deeply pipelined hardware CG implementation, targeted at modern FPGA architectures. This implementation is particularly suited for accelerating multiple small-to-medium-sized dense systems of linear equations and can be used as a stand-alone solver or as building block to solve higher-order systems. In this article it is shown that through parallelization it is possible to convert the computation time per iteration for an ordernmatrix fromΘ(n2) clock cycles on a microprocessor toΘ(n) on a FPGA. Through deep pipelining it is also possible to solve several problems in parallel and maximize both performance and efficiency. I/O requirements are shown to be scalable and convergent to a constant value with the increase of matrix order. Post place-and-route results on a readily available VirtexII-6000 demonstrate sustained performance of 5 GFlops, and results on a Virtex5-330 indicate sustained performance of 35 GFlops. A comparison with an optimized software implementation running on a high-end CPU demonstrate that this FPGA implementation represents a significant speedup of at least an order of magnitude.
登录
查看更多内容
DOI:
10.1007/978-3-540-78610-8_10
发表时间:
2008
期刊:
The American Mathematical Monthly
影响因子:
--
作者:
Antonio Roldao Lopes;G. Constantinides
通讯作者:
G. Constantinides
影响因子:
5.3
作者:
J. Kurzak;A. Buttari;J. Dongarra
通讯作者:
J. Dongarra
DOI:
--
发表时间:
2008
期刊:
International Conference on Field-Programmable Logic and Applications
影响因子:
--
作者:
M. Langhammer
通讯作者:
M. Langhammer
DOI:
--
发表时间:
2005
期刊:
影响因子:
--
作者:
I. Pournara;C. Bouganis;G. Constantinides
通讯作者:
G. Constantinides
DOI:
--
发表时间:
2005
期刊:
Parallel Processing and Applied Mathematics
影响因子:
--
作者:
O. Maslennikov;V. Lepekha;A. Sergyienko
通讯作者:
A. Sergyienko