A GPU Implementation of Computing Euclidean Distance Map with Efficient Memory Access

A GPU Implementation of Computing Euclidean Distance Map with Efficient Memory Access
复制标题

高效内存访问计算欧氏距离图的 GPU 实现

DOI:
--
复制
发表时间:
2011
期刊:
2011 Second International Conference on Networking and Computing
影响因子:
--
通讯作者:
K. Nakano
K. Nakano
中科院分区:
--
文献类型:
--
作者:
Duhu Man;K. Uda;Yasuaki Ito;K. Nakano

文献摘要

被引文献

相似文献

具有许多处理单元的最新图形处理单元(GPU)可用于通用并行计算。为了利用GPU强大的计算能力,GPU被广泛用于通用处理。由于GPU具有非常高的内存带宽,因此GPU的性能在很大程度上取决于内存访问。本文的主要贡献是提出了一个GPU实现计算欧几里德距离图(EDM)与高效的内存访问。给定一个2-D二值图像,EDM是一个大小相同的2-D数组,每个元素都存储到最近的黑色像素的欧几里得距离。在提出的GPU实现中,我们考虑了GPU系统的许多编程问题,如合并访问全局内存,共享内存库冲突和分区露营。在实践中,我们已经实现了我们的并行算法在以下两个现代GPU系统:特斯拉C1060和GTX 480,分别。实验结果表明,对于一幅大小为$9216的二值图像, imes 9216$,我们的实现可以实现一个加速因子为52的顺序算法实现。
Recent Graphics Processing Units (GPUs), which have many processing units, can be used for general purpose parallel computation. To utilize the powerful computing ability, GPUs are widely used for general purpose processing. Since GPUs have very high memory bandwidth, the performance of GPUs greatly depends on memory access. The main contribution of this paper is to present a GPU implementation of computing Euclidean Distance Map (EDM) with efficient memory access. Given a 2-D binary image, EDM is a 2-D array of the same size such that each element is storing the Euclidean distance to the nearest black pixel. In the proposed GPU implementation, we have considered many programming issues of the GPU system such as coalescing access of global memory, shared memory bank conflicts and partition camping. In practice, we have implemented our parallel algorithm in the following two modern GPU systems: Tesla C1060 and GTX 480, respectively. The experimental results have shown that, for an input binary image with size of $9216 imes 9216$, our implementation can achieve a speedup factor of 52 over the sequential algorithm implementation.