Efficient OpenCL Accelerators for Canny Edge Detection Algorithm on a CPU-FPGA Platform

Efficient OpenCL Accelerators for Canny Edge Detection Algorithm on a CPU-FPGA Platform
复制标题

DOI:
10.1109/reconfig48160.2019.8994769
复制
发表时间:
2019-12
期刊:
2019 International Conference on ReConFigurable Computing and FPGAs (ReConFig)
影响因子:
--
通讯作者:
Samah Rahamneh;L. Sawalha
Samah Rahamneh;L. Sawalha
中科院分区:
其他
文献类型:
--
作者:
Samah Rahamneh;L. Sawalha

文献摘要

相似文献

由于移动和边缘设备产生的海量数据,当前和新兴应用(如图像/视频处理)的处理需求正在增加。这给从智能手机到云和数据中心的各种计算系统带来了挑战。由于能够适应不同的工作负载需求,异质计算显示了其作为一种高效计算模型的能力。现场可编程门阵列(现场可编程门阵列)提供功率和性能优势,已被用于从嵌入式系统到云的许多应用领域。在本文中,我们使用了一种紧密耦合的CPU-FPGA异质系统来加速一种基于滑动窗口的图像处理算法--Canny边缘检测器。我们使用两种不同的实现加速了Canny:代码分区和数据分区。在数据分区的实现中,我们提出了一种基于加权轮询的算法,该算法对输入图像进行分区,并根据延迟在CPU和FPGA之间分配负载。文中还比较了所提出的加速器在单独的CPU和FPGA实现下的性能。使用基于CPU-FPGA的混合算法,我们在仅使用CPU的情况下实现了高达4.8倍的加速比,在仅使用FPGA的实现上实现了高达2.1倍的加速比。此外,我们的算法的估计总能量消耗比仅使用CPU的实现更高效。我们的结果显示,与仅使用CPU的实现相比,能量延迟乘积(EDP)显著降低,并且与仅使用FPGA的实现的EDP结果相当。
The processing demands of current and emerging applications, such as image/video processing, are increasing due to the deluge of data, generated by mobile and edge devices. This raises challenges for a vast range of computing systems, starting from smart-phones and reaching cloud and data centers. Heterogeneous computing demonstrates its ability as an efficient computing model due to its capability to adapt to various workload requirements. Field programmable gate arrays (FPGAs) provide power and performance benefits and have been used in many application domains from embedded systems to the cloud. In this paper, we used a closely coupled CPU-FPGA heterogeneous system to accelerate a sliding window based image processing algorithm, Canny edge detector. We accelerated Canny using two different implementations: Code partitioned and data partitioned. In the data partitioned implementation, we proposed a weighted round robin based algorithm that partitions input images and distributes the load between the CPU and the FPGA based on latency. The paper also compares the performance of the proposed accelerators with separate CPU and FPGA implementations. Using our hybrid CPU-FPGA based algorithm, we achieved a speedup up to 4.8× over a CPU-only and up to 2.1× over a FPGA-only implementations. Moreover, the estimated total energy consumption of our algorithm is more efficient than a CPU-only implementation. Our results show a significant reduction in energy delay product (EDP) compared to the CPU-only implementation, and comparable EDP results to the FPGA-only implementation.