Optimizing memory bandwidth use and performance for matrix-vector multiplication in iterative methods
Optimizing memory bandwidth use and performance for matrix-vector multiplication in iterative methods
复制标题
优化迭代方法中矩阵向量乘法的内存带宽使用和性能
DOI:
10.1145/2000832.2000834
复制
发表时间:
2011
影响因子:
2.3
通讯作者:
Boland D
中科院分区:
文献类型:
--
作者:
Boland D
Computing the solution to a system of linear equations is a fundamental problem in scientific computing, and its acceleration has drawn wide interest in the FPGA community [Morris et al. 2006; Zhang et al. 2008; Zhuo and Prasanna 2006]. One class of algorithms to solve these systems, iterative methods, has drawn particular interest, with recent literature showing large performance improvements over General-Purpose Processors (GPPs) [Lopes and Constantinides 2008]. In several iterative methods, this performance gain is largely a result of parallelization of the matrix-vector multiplication, an operation that occurs in many applications and hence has also been widely studied on FPGAs [Zhuo and Prasanna 2005; El-Kurdi et al. 2006]. However, whilst the performance of matrix-vector multiplication on FPGAs is generally I/O bound [Zhuo and Prasanna 2005], the nature of iterative methods allows the use of on-chip memory buffers to increase the bandwidth, providing the potential for significantly more parallelism [deLorimier and DeHon 2005]. Unfortunately, existing approaches have generally only either been capable of solving large matrices with limited improvement over GPPs [Zhuo and Prasanna 2005; El-Kurdi et al. 2006; deLorimier and DeHon 2005], or achieve high performance for relatively small matrices [Lopes and Constantinides 2008; Boland and Constantinides 2008]. This article proposes hardware designs to take advantage of symmetrical and banded matrix structure, as well as methods to optimize the RAM use, in order to both increase the performance and retain this performance for larger-order matrices.
登录
查看更多内容
DOI:
10.1016/b978-0-12-637475-9.50006-6
发表时间:
1988
期刊:
arXiv: Soft Condensed Matter
影响因子:
--
作者:
G. Sewell
通讯作者:
G. Sewell
DOI:
10.1007/978-3-540-78610-8_10
发表时间:
2008
期刊:
The American Mathematical Monthly
影响因子:
--
作者:
Antonio Roldao Lopes;G. Constantinides
通讯作者:
G. Constantinides
DOI:
--
发表时间:
2010
期刊:
International Workshop on Applied Reconfigurable Computing
影响因子:
--
作者:
D. Boland;G. Constantinides
通讯作者:
G. Constantinides
DOI:
10.1145/2133352.2133358
发表时间:
2008-12
期刊:
2008 International Conference on Field-Programmable Technology
影响因子:
--
作者:
Wei Zhang-;Vaughn Betz;Jonathan Rose
通讯作者:
Wei Zhang-;Vaughn Betz;Jonathan Rose
DOI:
--
发表时间:
2006
期刊:
2006 14th Annual IEEE Symposium on Field-Programmable Custom Computing Machines
影响因子:
--
作者:
G. R. Morris;V. Prasanna;Richard D. Anderson
通讯作者:
Richard D. Anderson