Accelerating Matrix Operations with Improved Deeply Pipelined Vector Reduction
Accelerating Matrix Operations with Improved Deeply Pipelined Vector Reduction
复制标题
DOI:
10.1109/tpds.2011.141
复制
发表时间:
2012-02
影响因子:
5.3
通讯作者:
Yi-Gang Tai;D. Lo;K. Psarris
中科院分区:
文献类型:
--
作者:
Yi-Gang Tai;D. Lo;K. Psarris
Many scientific or engineering applications involve matrix operations, in which reduction of vectors is a common operation. If the core operator of the reduction is deeply pipelined, which is usually the case, dependencies between the input data elements cause data hazards. To tackle this problem, we propose a new reduction method with low latency and high pipeline utilization. The performance of the proposed design is evaluated for both single data set and multiple data set scenarios. Further, QR decomposition is used to demonstrate how the proposed method can accelerate its execution. We implement the design on an FPGA and compare its results to other methods.