Effective Implementation of Edge-Preserving Filtering on CPU Microarchitectures

Effective Implementation of Edge-Preserving Filtering on CPU Microarchitectures
复制标题

DOI:
10.3390/app8101985
复制
发表时间:
2018-10
期刊:
影响因子:
--
通讯作者:
Y. Maeda;Norishige Fukushima;H. Matsuo
Y. Maeda;Norishige Fukushima;H. Matsuo
中科院分区:
--
文献类型:
--
作者:
Y. Maeda;Norishige Fukushima;H. Matsuo

文献摘要

被引文献

相似文献

在本文中,我们提出了用于边缘保护过滤的加速方法。过滤器本质上包括在IEEE标准754中定义的非规范化数字。不合规数的处理的计算成本高于正常数字。因此,边缘滤波的计算性能严重降低。我们提出了防止发生加速数量数字的方法的方法。此外,我们通过仔细处理内核重量来验证基于中央加工单元的微体系结构的变化,验证了边缘保护过滤的有效矢量化。实验结果表明,所提出的方法比直接实现双边滤波和非本地均值过滤的速度要快五倍,而过滤器保持高精度。此外,我们显示了每个中央处理单元微体系结构的有效矢量化。双侧过滤器的实施速度比OpenCV快14倍。提出的方法和矢量化对于实时任务(例如图像编辑)是实用的。
In this paper, we propose acceleration methods for edge-preserving filtering. The filters natively include denormalized numbers, which are defined in IEEE Standard 754. The processing of the denormalized numbers has a higher computational cost than normal numbers; thus, the computational performance of edge-preserving filtering is severely diminished. We propose approaches to prevent the occurrence of the denormalized numbers for acceleration. Moreover, we verify an effective vectorization of the edge-preserving filtering based on changes in microarchitectures of central processing units by carefully treating kernel weights. The experimental results show that the proposed methods are up to five-times faster than the straightforward implementation of bilateral filtering and non-local means filtering, while the filters maintain the high accuracy. In addition, we showed effective vectorization for each central processing unit microarchitecture. The implementation of the bilateral filter is up to 14-times faster than that of OpenCV. The proposed methods and the vectorization are practical for real-time tasks such as image editing.