Almost optimal column-wise prefix-sum computation on the GPU
Almost optimal column-wise prefix-sum computation on the GPU
复制标题
DOI:
10.1007/s11227-018-2242-8
复制
发表时间:
2017-09
期刊:
影响因子:
--
通讯作者:
Hiroki Tokura;Toru Fujita;K. Nakano;Yasuaki Ito;J. Bordim
中科院分区:
文献类型:
--
作者:
Hiroki Tokura;Toru Fujita;K. Nakano;Yasuaki Ito;J. Bordim
Row-wise and column-wise prefix-sum computation of a matrix has many applications in the area of image processing such as computation of the summed area table and the Euclidean distance map. It is known that the prefix-sums of a one-dimensional array can be computed efficiently on the GPU. Hence, row-wise prefix-sums of a matrix can also be computed efficiently on the GPU by executing this prefix-sum algorithm for every row in parallel. However, the same approach does not work well for computing column-wise prefix-sums due to inefficient stride memory access to the global memory is performed. The main contribution of this paper is to present an almost optimal column-wise prefix-sum algorithm on the GPU. Quite surprisingly, experimental results using NVIDIA TITAN X show that our column-wise prefix-sum algorithm runs only 2–6% slower than matrix duplication. Thus, our column-wise prefix-sum algorithm is almost optimal.