A Distributed Canny Edge Detector: Algorithm and FPGA Implementation

A Distributed Canny Edge Detector: Algorithm and FPGA Implementation
复制标题

DOI:
10.1109/tip.2014.2311656
复制
发表时间:
2014-03
影响因子:
10.6
通讯作者:
Qian Xu;Srenivas Varadarajan;C. Chakrabarti;Lina Karam
Qian Xu;Srenivas Varadarajan;C. Chakrabarti;Lina Karam
中科院分区:
计算机科学1区
文献类型:
--
作者:
Qian Xu;Srenivas Varadarajan;C. Chakrabarti;Lina Karam

文献摘要

被引文献

相似文献

Canny边缘检测算法以其优越的性能成为应用最广泛的边缘检测算法之一。不幸的是,与其他边缘检测算法相比,它不仅计算量更大,而且由于它基于帧级统计,因此延迟也更高。在本文中,我们提出了一种机制,在块级实现Canny算法,与原始帧级Canny算法相比,边缘检测性能没有任何损失。在块级直接应用原Canny算法,由于原Canny算法是基于帧级统计计算高、低阈值,导致光滑区域边缘过多,高细节区域显著边缘丢失。为了解决这一问题,我们提出了一种分布式Canny边缘检测算法,该算法根据图像块类型和梯度的局部分布自适应计算边缘检测阈值。此外,新算法使用非均匀梯度幅度直方图来计算基于块的滞后阈值。由此产生的基于块的算法具有显著降低的延迟,并且可以很容易地与其他基于块的图像编解码器集成。它能够支持高分辨率的图像和视频的快速边缘检测,包括全高清,因为延迟现在是块大小的函数,而不是帧大小。此外,定量一致性评估和主观测试表明,该算法的边缘检测性能优于原始的基于帧的算法,特别是当图像中存在噪声时。最后,采用32计算引擎架构实现了该算法,并在Xilinx Virtex-5 FPGA上进行了合成。在USC SIPI数据库中,当时钟频率为100 MHz时,该合成架构检测512 × 512图像的边缘只需要0.721 ms(包括SRAM读/写时间和计算时间),比现有的FPGA和GPU实现更快。
The Canny edge detector is one of the most widely used edge detection algorithms due to its superior performance. Unfortunately, not only is it computationally more intensive as compared with other edge detection algorithms, but it also has a higher latency because it is based on frame-level statistics. In this paper, we propose a mechanism to implement the Canny algorithm at the block level without any loss in edge detection performance compared with the original frame-level Canny algorithm. Directly applying the original Canny algorithm at the block-level leads to excessive edges in smooth regions and to loss of significant edges in high-detailed regions since the original Canny computes the high and low thresholds based on the frame-level statistics. To solve this problem, we present a distributed Canny edge detection algorithm that adaptively computes the edge detection thresholds based on the block type and the local distribution of the gradients in the image block. In addition, the new algorithm uses a nonuniform gradient magnitude histogram to compute block-based hysteresis thresholds. The resulting block-based algorithm has a significantly reduced latency and can be easily integrated with other block-based image codecs. It is capable of supporting fast edge detection of images and videos with high resolutions, including full-HD since the latency is now a function of the block size instead of the frame size. In addition, quantitative conformance evaluations and subjective tests show that the edge detection performance of the proposed algorithm is better than the original frame-based algorithm, especially when noise is present in the images. Finally, this algorithm is implemented using a 32 computing engine architecture and is synthesized on the Xilinx Virtex-5 FPGA. The synthesized architecture takes only 0.721 ms (including the SRAM READ/WRITE time and the computation time) to detect edges of 512 × 512 images in the USC SIPI database when clocked at 100 MHz and is faster than existing FPGA and GPU implementations.