FPGA-Based High-Performance and Scalable Block LU Decomposition Architecture

FPGA-Based High-Performance and Scalable Block LU Decomposition Architecture
复制标题

DOI:
10.1109/tc.2011.24
复制
发表时间:
2012
影响因子:
3.7
通讯作者:
M. Jaiswal;N. Chandrachoodan
M. Jaiswal;N. Chandrachoodan
中科院分区:
计算机科学2区
文献类型:
--
作者:
M. Jaiswal;N. Chandrachoodan

文献摘要

被引文献

相似文献

将矩阵分解为上、下三角矩阵(LU分解)是许多科学和工程应用的重要组成部分,而块LU分解算法是一种非常适合并行硬件实现的方法。本文提出了一种利用FPGA硬件加速实现块LU分解算法的方法。与文献中报道的大多数以前的方法不同,该方法不假设矩阵可以完全存储在芯片上。研究了不同的现场可编程门阵列结构下的存储器访问,并给出了扩展良好的操作时间表。该设计已经针对FPGA目标进行了综合,并且可以很容易地进行重定向。该设计超越了以前的硬件实现,以及包括工作站上的ATLAS和MKL库在内的已调优软件实现。
Decomposition of a matrix into lower and upper triangular matrices (LU decomposition) is a vital part of many scientific and engineering applications, and the block LU decomposition algorithm is an approach well suited to parallel hardware implementation. This paper presents an approach to speed up implementation of the block LU decomposition algorithm using FPGA hardware. Unlike most previous approaches reported in the literature, the approach does not assume the matrix can be stored entirely on chip. The memory accesses are studied for various FPGA configurations, and a schedule of operations for scaling well is shown. The design has been synthesized for FPGA targets and can be easily retargeted. The design outperforms previous hardware implementations, as well as tuned software implementations including the ATLAS and MKL libraries on workstations.