FPGA-based circular hough transform with graph clustering for vision-based multi-robot tracking

FPGA-based circular hough transform with graph clustering for vision-based multi-robot tracking
复制标题

基于 FPGA 的圆形霍夫变换和图形聚类,用于基于视觉的多机器人跟踪

DOI:
10.1109/reconfig.2015.7393313
复制
发表时间:
2015
期刊:
2015 International Conference on ReConFigurable Computing and FPGAs (ReConFig)
影响因子:
--
通讯作者:
U. Rückert
U. Rückert
中科院分区:
--
文献类型:
--
作者:
Arif Irwansyah;O. Ibraheem;J. Hagemeyer;Mario Porrmann;U. Rückert

文献摘要

参考文献

被引文献

相似文献

基于形状的对象检测和识别是在计算机视觉领域中经常使用的方法。圆形检测的众所周知的算法是圆形霍夫变换(CHT)。这种霍夫变换算法需要巨大的内存空间和大量的计算资源。可以使用现场可编程栅极阵列(FPGA)的硬件加速器来有效处理此类计算密集型应用程序。在本文中,我们为CHT算法提供了基于资源有效的基于FPGA的体系结构。此外,我们通过将CHT算法与图形聚类相结合,从而引入了独特的方法。这些算法的组合及其在Xilinx virtex-4 FPGA上的实现用于支持基于实时视觉的多机器人跟踪。此外,提出了有效的体系结构,以显着减少CHT模块中所需的内存。对于图形聚类模块,实现了无乘数距离计算单元,从而大大降低了所需的FPGA资源。所提出的CHT设计可以处理多机器人定位,精度为97%,支持1024x1024的最大视频分辨率,每秒128帧,导致134 mpixel/s。与嵌入式处理器,FPGA和通用CPU上的其他实现相比,我们的设计提供了明显更高的吞吐量。与3.2 GHz桌面CPU上的OPENCV实施相比,我们的实施实现了超过5.7的速度。
Shape-based object detection and recognition are frequently used methods in the field of computer vision. A well-known algorithm for circle detection is the Circular Hough Transform (CHT). This Hough Transform algorithm needs a huge memory space and large computational resources. Field Programmable Gate Array (FPGA)-based hardware accelerators can be used to efficiently handle such compute-intensive applications. In this paper, we present a resource-efficient FPGA-based architecture for the CHT algorithm. Additionally, we introduce a unique approach by combining the CHT algorithm with graph clustering. The combination of these algorithms and their implementation on a Xilinx Virtex-4 FPGA is used to support real-time vision-based multi-robot tracking. Furthermore, an efficient architecture is proposed to significantly reduce the required memory in the CHT module. For the Graph Clustering module, a multiplier-less distance calculation unit is implemented, significantly reducing the required FPGA resources. The proposed CHT design can handle multi-robot localization with an accuracy of 97 %, supporting a maximum video resolution of 1024x1024 with 128 frames per second, resulting in 134 MPixel/s. Our design provides significantly higher throughput compared to other implementations on embedded processors, FPGAs, and general purpose CPUs. Compared to an OpenCV implementation on a 3.2 GHz desktop CPU, our implementation achieves a speed- up of more than 5.7.
DOI: 10.1145/361237.361242
发表时间: 1972-01-01
影响因子: 22.7
作者:
DUDA, RO;HART, PE
通讯作者: HART, PE