Optimising Memory Bandwidth Use for Matrix-Vector Multiplication in Iterative Methods

Optimising Memory Bandwidth Use for Matrix-Vector Multiplication in Iterative Methods
复制标题

优化迭代方法中矩阵向量乘法的内存带宽使用

DOI:
--
复制
发表时间:
2010
期刊:
International Workshop on Applied Reconfigurable Computing
影响因子:
--
通讯作者:
G. Constantinides
G. Constantinides
中科院分区:
--
文献类型:
--
作者:
D. Boland;G. Constantinides

文献摘要

被引文献

相似文献

计算线性方程组的解是科学计算中的一个基本问题,其加速在FPGA社区中引起了广泛的兴趣[1,2,3]。一类算法来解决这些系统,迭代方法,引起了特别的兴趣,最近的文献显示出大的性能改进,通用处理器(GPP)。在几种迭代方法中,这种性能增益在很大程度上是矩阵向量乘法的并行化的结果,矩阵向量乘法是一种在许多应用中发生的操作,因此也在FPGA上得到了广泛的研究[4,5]。然而,虽然FPGA上的矩阵向量乘法的性能通常受I/O限制[4],但迭代方法的性质允许使用片内存储器缓冲器来增加带宽,从而提供了显著更多并行性的潜力[6]。不幸的是,现有的方法通常只能解决大型矩阵,对GPP [4,5,6]的改进有限,或者对相对较小的矩阵实现高性能[2,3]。本文提出了硬件设计,以利用对称和带状矩阵结构,以及方法来优化RAM的使用,以提高性能,并保持这种性能为较大的顺序矩阵。
Computing the solution to a system of linear equations is a fundamental problem in scientific computing, and its acceleration has drawn wide interest in the FPGA community [1, 2, 3]. One class of algorithms to solve these systems, iterative methods, has drawn particular interest, with recent literature showing large performance improvements over general purpose processors (GPPs). In several iterative methods, this performance gain is largely a result of parallelisation of the matrixvector multiplication, an operation that occurs in many applications and hence has also been widely studied on FPGAs [4, 5]. However, whilst the performance of matrix-vector multiplication on FPGAs is generally I/O bound [4], the nature of iterative methods allows the use of onchip memory buffers to increase the bandwidth, providing the potential for significantly more parallelism [6]. Unfortunately, existing approaches have generally only either been capable of solving large matrices with limited improvement over GPPs [4,5,6], or achieve high performance for relatively small matrices [2,3]. This paper proposes hardware designs to take advantage of symmetrical and banded matrix structure, as well as methods to optimise the RAM use, in order to both increase the performance and retain this performance for larger order matrices.