Efficient Computational Scheduling of Box and Gaussian FIR Filtering for CPU Microarchitecture

Efficient Computational Scheduling of Box and Gaussian FIR Filtering for CPU Microarchitecture
复制标题

DOI:
10.23919/apsipa.2018.8659674
复制
发表时间:
2018-11
期刊:
2018 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
影响因子:
--
通讯作者:
Norishige Fukushima;Y. Maeda;Yuki Kawasaki;Masahiro Nakamura;Tomoaki Tsumura;Kenjiro Sugimoto;S. Kamata
Norishige Fukushima;Y. Maeda;Yuki Kawasaki;Masahiro Nakamura;Tomoaki Tsumura;Kenjiro Sugimoto;S. Kamata
中科院分区:
其他
文献类型:
--
作者:
Norishige Fukushima;Y. Maeda;Yuki Kawasaki;Masahiro Nakamura;Tomoaki Tsumura;Kenjiro Sugimoto;S. Kamata

文献摘要

相似文献

在本文中,我们提出了有效的计算调度盒和高斯滤波。这些滤波器是基本工具,用于各种应用。这些FIR滤波器的简单实现的计算阶数为$O(r^{2})$,其中$r$是核半径。一个可分离的实现将阶数减少到$O(r)$,但需要两次过滤。一个递归的表示法会将顺序大大地分解为O(1)$,但也需要两次或更多次过滤。有效的表示减少了算术运算的数量,然而,数据I/O对计算时间的影响变得占主导地位。本文对$O(1)$盒滤波器和高斯滤波器的计算调度进行了优化,以充分利用高速缓冲存储器来减少数据I/O的计算时间。实验结果表明,该调度算法具有更高的计算性能比传统的实现。
In this paper, we propose efficient computational scheduling of box and Gaussian filtering. These filters are fundamental tools and used for various applications. The computational order of the naïve implementations of these FIR filters are $O(r^{2})$, where $r$ is the kernel radius. A separable implementation reduces the order into $O(r)$ but requires twice times of filtering. A recursive representation dramatically sheds the order into $O(1)$ but also needs twice or more times filtering. The efficient representation curtails the number of arithmetic operations; however, the influence of data I/O for the computational time becomes dominant. In this paper, we optimize the computational scheduling of $O(1)$ box and Gaussian filters to competently utilize cache memory for reducing the computational time of data I/O. Experimental results show that the proposed scheduling has higher computational performance than the conventional implementation.