A High Performance and Memory Efficient LU Decomposer on FPGAs

A High Performance and Memory Efficient LU Decomposer on FPGAs
复制标题

FPGA 上的高性能和内存高效的 LU 分解器

DOI:
10.1109/tc.2010.278
复制
发表时间:
2012-03
影响因子:
3.7
通讯作者:
Peterson, Gregory D.
Peterson, Gregory D.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Wu, Guiming;Dou, Yong;Sun, Junqing;Peterson, Gregory D.

文献摘要

参考文献

被引文献

相似文献

稠密矩阵的LU分解是一种重要的线性代数核,在科学和工程应用中有着广泛的应用。为了在FPGA上高效地实现大规模矩阵LU分解,提出了一种适用于任意矩阵大小的FPGA分块LU分解算法。我们的算法适用于一系列的转换,包括循环阻塞和时空映射,到顺序非阻塞LU分解。我们还介绍了一个高性能和内存效率的硬件架构,它主要由一个线性阵列的处理元件(PE),实现我们的块LU分解算法。我们的设计可以在各种硬件资源限制下达到最佳性能。此外,我们的算法和设计可以很容易地扩展到多FPGA平台,通过使用块循环数据分配和FPGA间的通信方案。Xilinx Virtex-5 XC 5VLX 330 FPGA在我们自行设计的PCI-Express卡上可以集成36个PE,在133 MHz时达到8.50 GFLOPS的持续性能,矩阵大小为16,384,优于几个通用处理器。对于Xilinx Virtex-6 XC 6VLX 760(一种较新的FPGA),我们预测总共可以集成180个PE,在200 MHz时达到70.66 GFLOPS。与以前的工作相比,我们的设计可以集成到同一个FPGA的PE数量的两倍,并具有显着更高的性能。
LU decomposition for dense matrices is an important linear algebra kernel that is widely used in both scientific and engineering applications. To efficiently perform large matrix LU decomposition on FPGAs with limited local memory, a block LU decomposition algorithm on FPGAs applicable to arbitrary matrix size is proposed. Our algorithm applies a series of transformations, including loop blocking and space-time mapping, onto sequential nonblocking LU decomposition. We also introduce a high performance and memory efficient hardware architecture, which mainly consists of a linear array of processing elements (PEs), to implement our block LU decomposition algorithm. Our design can achieve optimum performance under various hardware resource constraints. Furthermore, our algorithm and design can be easily extended to the multi-FPGA platform by using a block-cyclic data distribution and inter-FPGA communication scheme. A total of 36 PEs can be integrated into a Xilinx Virtex-5 XC5VLX330 FPGA on our self-designed PCI-Express card, reaching a sustained performance of 8.50 GFLOPS at 133 MHz for a matrix size of 16,384, which outperforms several general-purpose processors. For a Xilinx Virtex-6 XC6VLX760, a newer FPGA, we predict that a total of 180 PEs can be integrated, reaching 70.66 GFLOPS at 200 MHz. Compared to the previous work, our design can integrate twice the number of PEs into the same FPGA and has significantly higher performance.
DOI: 10.1109/ipdps.2004.1303134
发表时间: 2004-04
期刊: 18th International Parallel and Distributed Processing Symposium, 2004. Proceedings.
影响因子: --
作者:
G. Govindu;S. Choi;V. Prasanna;V. Daga;Sridhar Gangadharpalli;V. Sridhar
通讯作者: G. Govindu;S. Choi;V. Prasanna;V. Daga;Sridhar Gangadharpalli;V. Sridhar
DOI: --
发表时间: 2008
期刊: --
影响因子: --
作者:
通讯作者: --
DOI: 10.1137/1026003
发表时间: 1984
期刊: Siam Review
影响因子: 10.2
作者:
J. Dongarra;F. Gustavson;A. Karp
通讯作者: J. Dongarra;F. Gustavson;A. Karp
DOI: 10.1007/978-0-387-09766-4_2177
发表时间: 2011
期刊: --
影响因子: --
作者:
Rajesh K. Karmani;G. Agha;M. Squillante;J. Seiferas;M. Brezina;Jonathan Hu;R. Tuminaro;P. Sanders;Jesper Larsson Träffe;R. Geijn;J. Träff;MS Benjamin Sander;J. Gustafson;R. Dror;C. Young;D. Shaw;Calvin Lin;Jenq-Kuen Lee;Rong-Guey Chang;Chi-Bang Kuan;G. Kollias;A. Grama;Zhiyuan Li;R. Clint Whaley;R. Vuduc
通讯作者: Rajesh K. Karmani;G. Agha;M. Squillante;J. Seiferas;M. Brezina;Jonathan Hu;R. Tuminaro;P. Sanders;Jesper Larsson Träffe;R. Geijn;J. Träff;MS Benjamin Sander;J. Gustafson;R. Dror;C. Young;D. Shaw;Calvin Lin;Jenq-Kuen Lee;Rong-Guey Chang;Chi-Bang Kuan;G. Kollias;A. Grama;Zhiyuan Li;R. Clint Whaley;R. Vuduc
DOI: 10.1109/fpl.2006.311240
发表时间: 2006-08
期刊: 2006 International Conference on Field Programmable Logic and Applications
影响因子: --
作者:
K. Turkington;K. Masselos;G. Constantinides;P. Leong
通讯作者: K. Turkington;K. Masselos;G. Constantinides;P. Leong