Video Coding With Rate-Distortion Optimized Transform

Video Coding With Rate-Distortion Optimized Transform
复制标题

DOI:
10.1109/tcsvt.2011.2158363
复制
发表时间:
2012
影响因子:
8.4
通讯作者:
Xin Zhao;Li Zhang;Siwei Ma;Wen Gao
Xin Zhao;Li Zhang;Siwei Ma;Wen Gao
中科院分区:
工程技术1区
文献类型:
--
作者:
Xin Zhao;Li Zhang;Siwei Ma;Wen Gao

文献摘要

被引文献

相似文献

基于块的离散余弦变换(DCT)已成功应用于多个国际图像/视频编码标准,例如MPEG-2、H.264/AVC,因为它可以在性能和复杂性之间实现良好的权衡。尽管 DCT 理论上近似一阶马尔可夫条件下的最佳 Karhunen-Loève 变换,但由于视频内容的非平稳性质,一组固定的变换基函数 (TBF) 无法有效处理所有情况。为了进一步提高基于块的变换编码的性能,在本文中,我们提出了速率失真优化变换(RDOT)的设计,它有助于帧内和帧间编码。 RDOT 与传统 DCT 之间最重要的区别在于,在所提出的方法中,变换是通过从离线训练获得的多个 TBF 候选来实现的。利用这一特性,对于每个残差块的编码,编码器能够在率失真性能方面选择最佳的TBF集合,并且在变换域中实现更好的能量压缩。为了获得最佳的候选 TBF 组,我们开发了一种用于离线训练的两步迭代优化技术,在每次迭代中对 TBF 候选进行细化,直到训练过程收敛。此外,本文还对候选TBF的最优组进行了分析,并详细描述了所提出算法在最新VCEG关键技术领域软件平台上的实际实现。大量的实验结果表明,与最先进的 H.264/AVC 视频编码标准中采用的传统基于 DCT 的变换方案相比,我们提出的方法在帧内和帧间编码方面都实现了编码性能的显着提高。
Block-based discrete cosine transform (DCT) has been successfully adopted into several international image/video coding standards, e.g., MPEG-2, H.264/AVC, as it can achieve a good tradeoff between performance and complexity. Although DCT theoretically approximates the optimum Karhunen-Loève transform under first-order Markov conditions, one fixed set of transform basis functions (TBF) cannot handle all the cases efficiently due to the non-stationary nature of video contents. To further improve the performance of block-based transform coding, in this paper, we present the design of rate-distortion optimized transform (RDOT) which contributes to both intraframe and interframe coding. The most important property which makes a difference between RDOT and the conventional DCT is that, in the proposed method, transform is implemented with multiple TBF candidates which are obtained from off-line training. With this feature, for coding each residual block, the encoder is capable to select the optimal set of TBF in terms of rate-distortion performance, and better energy compaction is achieved in the transform domain. To obtain an optimum group of candidate TBF, we have developed a two-step iterative optimization technique for the off-line training, with which the TBF candidates are refined at each iteration until the training process becomes converged. Moreover, analysis on the optimal group of candidate TBF is also presented in this paper, with a detailed description of a practical implementation for the proposed algorithm on the latest VCEG key technical area software platform. Extensive experimental results show that, compared with the conventional DCT-based transform scheme adopted into the state-of-the-art H.264/AVC video coding standard, significant improvement of coding performance has been achieved for both intraframe and interframe coding with our proposed method.