Memory Footprint Reduction for Power-Efficient Realization of 2-D Finite Impulse Response Filters

Memory Footprint Reduction for Power-Efficient Realization of 2-D Finite Impulse Response Filters
复制标题

DOI:
10.1109/tcsi.2013.2265953
复制
发表时间:
2014
期刊:
IEEE Transactions on Circuits and Systems I: Regular Papers
影响因子:
--
通讯作者:
B. K. Mohanty;P. Meher;S. Al-Maadeed;A. Amira
B. K. Mohanty;P. Meher;S. Al-Maadeed;A. Amira
中科院分区:
其他
文献类型:
--
作者:
B. K. Mohanty;P. Meher;S. Al-Maadeed;A. Amira

文献摘要

被引文献

相似文献

我们分析了内存占用和组合复杂性,得出了一种系统设计策略,为二维 (2-D) 有限脉冲响应 (FIR) 滤波器导出面积延迟功率高效的架构。我们为可分离和不可分离滤波器提出了新颖的基于块的结构,通过内存共享和内存重用以及适当的计算调度和存储架构设计来减少内存占用。与现有结构相比,所提出的结构的每个输出存储量 (SPO) 减少了 L 倍,每个输出能耗 (EPO) 减少了近 L 倍,其中 L 是输入块大小。它们涉及的算术资源是相应现有结构中最好的结构的 L 倍,并且在内存带宽 (MBW) 方面比其他结构产生的吞吐量高出 L 倍。我们还提出了可分离和不可分离滤波器组的单独通用结构,以及构成对称滤波器和通用滤波器的滤波器组的统一结构。与现有类似的统一结构相比,所提出的 6 个并行滤波器的统一结构涉及近 3.6L 倍的乘法器、3L 倍的加法器、更少的 (N2-N+2) 个寄存器,并且每个周期计算的滤波器输出量增加了 6L 倍,MBW 比现有设计少 6L 倍,其中 N 是每个维度的 FIR 滤波器大小。 ASIC综合结果表明,对于滤波器尺寸(4×4)、输入块尺寸L=4和图像尺寸(512×512),所提出的基于块的不可分离和通用不可分离结构分别比相应的现有结构少5.95倍和11.25倍的面积延迟积(ADP),以及5.81倍和15.63倍的EPO。与相应的现有结构相比,所提出的统一结构涉及的 ADP 减少了 4.64 倍,EPO 减少了 9.78 倍。
We have analyzed memory footprint and combinational complexity to arrive at a systematic design strategy to derive area-delay-power-efficient architectures for two-dimensional (2-D) finite impulse response (FIR) filter. We have presented novel block-based structures for separable and non-separable filters with less memory footprint by memory sharing and memory-reuse along with appropriate scheduling of computations and design of storage architecture. The proposed structures involve L times less storage per output (SPO), and nearly L times less energy consumption per output (EPO) compared with the existing structures, where L is the input block-size. They involve L times more arithmetic resources than the best of the corresponding existing structures, and produce L times more throughput with less memory band-width (MBW) than others. We have also proposed separate generic structures for separable and non-separable filter-banks, and a unified structure of filter-bank constituting symmetric and general filters. The proposed unified structure for 6 parallel filters involves nearly 3.6L times more multipliers, 3L times more adders, (N2-N+2) less registers than similar existing unified structure, and computes 6L times more filter outputs per cycle with 6L times less MBW than the existing design, where N is FIR filter size in each dimension. ASIC synthesis result shows that for filter size (4 × 4), input-block size L=4, and image-size (512 × 512), proposed block-based non-separable and generic non-separable structures, respectively, involve 5.95 times and 11.25 times less area-delay-product (ADP), and 5.81 times and 15.63 times less EPO than the corresponding existing structures. The proposed unified structure involves 4.64 times less ADP and 9.78 times less EPO than the corresponding existing structure.