Optimizing memory bandwidth use and performance for matrix-vector multiplication in iterative methods

Optimizing memory bandwidth use and performance for matrix-vector multiplication in iterative methods
复制标题

优化迭代方法中矩阵向量乘法的内存带宽使用和性能

DOI:
10.1145/2000832.2000834
复制
发表时间:
2011
影响因子:
2.3
通讯作者:
Boland D
Boland D
中科院分区:
计算机科学3区
文献类型:
--
作者:
Boland D

文献摘要

参考文献

被引文献

相似文献

计算线性方程组的解是科学计算中的一个基本问题,其加速引起了FPGA社区的广泛兴趣[Morris等人]。2006年;张等人。2008年;卓和普拉萨纳,2006年]。解决这些系统的一类算法,迭代方法,引起了特别的兴趣,最近的文献显示,与通用处理器(GPP)相比,性能有了很大的改进[Lope and Constantinides 2008]。在几种迭代方法中,这种性能提高在很大程度上是矩阵-向量乘法并行化的结果,这是一种在许多应用中发生的操作,因此也在FPGA上得到了广泛的研究[Choo和Prasanna 2005;El-Kurdi等人。2006年)。然而,虽然在现场可编程门阵列上矩阵-向量乘法的性能通常是I/O受限的[卓和普拉萨纳2005],但迭代方法的本质允许使用片上存储缓冲区来增加带宽,从而提供了显著增加并行度的可能性[Delorimier和DeHon2005]。不幸的是,现有的方法通常只能够求解大型矩阵,与GPP相比改进有限[Zuo和Prasanna 2005;El-Kurdi等人。2006年;Delorimier和DeHon2005年),或者在相对较小的矩阵上实现高性能[洛佩斯和康斯坦丁尼德2008年;博兰和康斯坦丁尼德2008年]。本文提出了利用对称和带状矩阵结构的硬件设计,以及优化RAM使用的方法,以便在高阶矩阵的情况下既提高性能又保持这一性能。
Computing the solution to a system of linear equations is a fundamental problem in scientific computing, and its acceleration has drawn wide interest in the FPGA community [Morris et al. 2006; Zhang et al. 2008; Zhuo and Prasanna 2006]. One class of algorithms to solve these systems, iterative methods, has drawn particular interest, with recent literature showing large performance improvements over General-Purpose Processors (GPPs) [Lopes and Constantinides 2008]. In several iterative methods, this performance gain is largely a result of parallelization of the matrix-vector multiplication, an operation that occurs in many applications and hence has also been widely studied on FPGAs [Zhuo and Prasanna 2005; El-Kurdi et al. 2006]. However, whilst the performance of matrix-vector multiplication on FPGAs is generally I/O bound [Zhuo and Prasanna 2005], the nature of iterative methods allows the use of on-chip memory buffers to increase the bandwidth, providing the potential for significantly more parallelism [deLorimier and DeHon 2005]. Unfortunately, existing approaches have generally only either been capable of solving large matrices with limited improvement over GPPs [Zhuo and Prasanna 2005; El-Kurdi et al. 2006; deLorimier and DeHon 2005], or achieve high performance for relatively small matrices [Lopes and Constantinides 2008; Boland and Constantinides 2008]. This article proposes hardware designs to take advantage of symmetrical and banded matrix structure, as well as methods to optimize the RAM use, in order to both increase the performance and retain this performance for larger-order matrices.
DOI: 10.1016/b978-0-12-637475-9.50006-6
发表时间: 1988
期刊: arXiv: Soft Condensed Matter
影响因子: --
作者:
G. Sewell
通讯作者: G. Sewell
DOI: 10.1007/978-3-540-78610-8_10
发表时间: 2008
期刊: The American Mathematical Monthly
影响因子: --
作者:
Antonio Roldao Lopes;G. Constantinides
通讯作者: G. Constantinides
优化迭代方法中矩阵向量乘法的内存带宽使用
DOI: --
发表时间: 2010
期刊: International Workshop on Applied Reconfigurable Computing
影响因子: --
作者:
D. Boland;G. Constantinides
通讯作者: G. Constantinides
DOI: 10.1145/2133352.2133358
发表时间: 2008-12
期刊: 2008 International Conference on Field-Programmable Technology
影响因子: --
作者:
Wei Zhang-;Vaughn Betz;Jonathan Rose
通讯作者: Wei Zhang-;Vaughn Betz;Jonathan Rose
将共轭梯度映射到 FPGA 增强可重构超级计算机的混合方法
DOI: --
发表时间: 2006
期刊: 2006 14th Annual IEEE Symposium on Field-Programmable Custom Computing Machines
影响因子: --
作者:
G. R. Morris;V. Prasanna;Richard D. Anderson
通讯作者: Richard D. Anderson