A 48 Cycles/MB H.264/AVC Deblocking Filter Architecture for Ultra High Definition Applications

A 48 Cycles/MB H.264/AVC Deblocking Filter Architecture for Ultra High Definition Applications
复制标题

DOI:
10.1587/transfun.e92.a.3203
复制
发表时间:
2009-12
期刊:
IEICE Trans. Fundam. Electron. Commun. Comput. Sci.
影响因子:
--
通讯作者:
Dajiang Zhou;Jinjia Zhou;Jiayi Zhu;Satoshi Goto
Dajiang Zhou;Jinjia Zhou;Jiayi Zhu;Satoshi Goto
中科院分区:
其他
文献类型:
--
作者:
Dajiang Zhou;Jinjia Zhou;Jiayi Zhu;Satoshi Goto

文献摘要

被引文献

相似文献

本文提出了一种H.264/AVC的高度并行的去块滤波器架构,该架构可以在48个时钟周期内处理一个宏块,并对小于100MHz的QFHD@60fps序列提供实时支持。在该架构中采用了分为两组的4个边缘滤波器,用于同时处理垂直和水平边缘,以提高其吞吐量。在并行度提高的同时,由于边缘滤波器的延迟性和分块算法的数据依赖性,产生了管道危险。为了解决这一问题,提出了一种消除管道气泡的锯齿形加工方案。然后根据处理进度导出体系结构的数据路径,并通过数据流合并进行优化,使逻辑和内部缓冲区的开销最小。同时,该架构的数据输入速率被设计为与其吞吐量相同,而输入数据的传输顺序也可以匹配锯齿形处理时间表。因此,为了速度匹配或数据重新排序,在去块滤波器和之前的组件之间不需要通信缓冲区。因此,在本设计中只需要一个24×64双端口SRAM作为内部缓冲区。当采用中芯国际130nm工艺合成时,该架构的栅极数为30.2万,考虑到其高性能,具有竞争力。
In this paper, a highly parallel deblocking filter architecture for H.264/AVC is proposed to process one macroblock in 48 clock cycles and give real-time support to QFHD@60fps sequences at less than 100MHz. 4 edge filters organized in 2 groups for simultaneously processing vertical and horizontal edges are applied in this architecture to enhance its throughput. While parallelism increases, pipeline hazards arise owing to the latency of edge filters and data dependency of deblocking algorithm. To solve this problem, a zig-zag processing schedule is proposed to eliminate the pipeline bubbles. Data path of the architecture is then derived according to the processing schedule and optimized through data flow merging, so as to minimize the cost of logic and internal buffer. Meanwhile, the architecture's data input rate is designed to be identical to its throughput, while the transmission order of input data can also match the zig-zag processing schedule. Therefore no intercommunication buffer is required between the deblocking filter and its previous component for speed matching or data reordering. As a result, only one 24×64 two-port SRAM as internal buffer is required in this design. When synthesized with SMIC 130nm process, the architecture costs a gate count of 30.2k, which is competitive considering its high performance.