Architecture Design for H.264/AVC Integer Motion Estimation with Minimum Memory Bandwidth

Architecture Design for H.264/AVC Integer Motion Estimation with Minimum Memory Bandwidth
复制标题

DOI:
10.1109/tce.2007.4341585
复制
发表时间:
2007-08
影响因子:
4.3
通讯作者:
Dongxiao Li;Wei Zheng;Ming Zhang
Dongxiao Li;Wei Zheng;Ming Zhang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Dongxiao Li;Wei Zheng;Ming Zhang

文献摘要

被引文献

相似文献

运动估计是视频编码系统中最关键的组成部分,也是影响计算复杂度和存储带宽的主要因素。针对H.264/AVC整数运动估计(IME),提出了一种新的存储器访问和计算高效的全搜索块匹配硬件结构。通过最高级别的片上数据重用,实现了对片外参考像素的一次访问,从而最小化了片外存储器带宽。通过分布式数据缓存和参考图像边界的虚拟连接,使得数据流量调度简单、规则、高效。计算引擎采用二维脉动处理器阵列以单指令多数据流(SIMD)方式计算绝对差值,并使用二维加法器树求和绝对差值,利用率均为100%。该结构完全支持H.264/AVC的可变块大小匹配,并且每周期一个搜索点可以产生41个绝对差值和(SADS),并且没有气泡。给出了该体系结构的参数化设计,并给出了标清数字电视编码应用的实现。理论分析和实验结果表明,该结构可以实现最小的片外存储带宽和最大的计算性能。
Motion estimation (ME) is the most critical component of a video coding system, and it also dominates the major part of computation complexity and memory bandwidth. For H.264/AVC integer motion estimation (IME), this paper presents a novel memory-access and computation efficient full-search block-matching hardware architecture. With the highest level of on-chip data reuse, one-access for off-chip reference pixels is achieved, and the off-chip memory bandwidth is thus minimized. By distributed data caching and virtual connection of reference picture boundaries, the data traffic scheduling is simple, regular and efficient. The computation engine employs a two-dimensional (2-D) systolic processor array to calculate the absolute differences in single-instruction multiple-data (SIMD) manner, and 2-D adder trees to sum up the absolute differences, all with 100% utilization. The proposed architecture fully supports variable block-size matching of H.264/AVC, and can produce 41 sums of absolute differences (SADs) for one search point every cycle without bubble. The architecture is described in parameterized design, and an implementation for standard-definition digital TV encoding applications is presented. Theoretical analysis and experimental results show that, the proposed architecture can achieve the minimum off-chip memory bandwidth and the maximum computational performance.