Accelerating Matrix Operations with Improved Deeply Pipelined Vector Reduction

Accelerating Matrix Operations with Improved Deeply Pipelined Vector Reduction
复制标题

DOI:
10.1109/tpds.2011.141
复制
发表时间:
2012-02
影响因子:
5.3
通讯作者:
Yi-Gang Tai;D. Lo;K. Psarris
Yi-Gang Tai;D. Lo;K. Psarris
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yi-Gang Tai;D. Lo;K. Psarris

文献摘要

被引文献

相似文献

许多科学或工程应用都涉及矩阵运算,其中向量归约是一种常见的运算。如果归约的核心运算符是深度流水线的(通常是这种情况),则输入数据元素之间的依赖性会导致数据冒险。为了解决这个问题,我们提出了一种新的减少方法,具有低延迟和高流水线利用率。所提出的设计的性能进行评估,为单一的数据集和多个数据集的情况下。此外,QR分解是用来证明如何所提出的方法可以加速其执行。我们在FPGA上实现了该设计,并将其结果与其他方法进行了比较。
Many scientific or engineering applications involve matrix operations, in which reduction of vectors is a common operation. If the core operator of the reduction is deeply pipelined, which is usually the case, dependencies between the input data elements cause data hazards. To tackle this problem, we propose a new reduction method with low latency and high pipeline utilization. The performance of the proposed design is evaluated for both single data set and multiple data set scenarios. Further, QR decomposition is used to demonstrate how the proposed method can accelerate its execution. We implement the design on an FPGA and compare its results to other methods.