Deep Learning-Based Intra Mode Derivation for Versatile Video Coding

Deep Learning-Based Intra Mode Derivation for Versatile Video Coding
复制标题

基于深度学习的多功能视频编码的帧内模式推导

DOI:
10.1145/3563699
复制
发表时间:
2022
期刊:
ACM Transactions on Multimedia Computing, Communications, and Applications
影响因子:
--
通讯作者:
Sam Kwong
Sam Kwong
中科院分区:
其他
文献类型:
--
作者:
Linwei Zhu;Yun Zhang;Na Li;Gangyi Jiang;Sam Kwong

文献摘要

相似文献

在码内编码中,通过率失真优化(RDO)从预定义的候选列表中获得最优的码内模式。除了消耗大量编码位的剩余信号外,还需要对最优的内模式进行编码并传输到解码器侧。为了进一步提高通用视频编码(VVC)中帧内编码的性能,本文提出了一种基于深度学习的帧内编码方法(DLIMD)。具体而言,将模内推导过程制定为一个多类分类任务,旨在跳过模内信令模块进行编码位缩减。DLIMD的结构是为了适应不同的量化参数设置和不同的编码块,包括非平方编码块,其中只需要一个训练模型。与现有的基于深度学习的分类问题不同的是,除了从特征学习网络中学习到的特征外,还将手工制作的特征输入到模式内衍生网络中。为了与传统方法竞争,在视频编解码器中使用一个附加的二进制标志来表示RDO所选择的方案。大量实验结果表明,该方法在VVC测试模型平台上对Y、U、V分量的码率平均降低2.28%、1.74%、2.18%,优于目前的研究成果。
In intra coding, Rate Distortion Optimization (RDO) is performed to achieve the optimal intra mode from a pre-defined candidate list. The optimal intra mode is also required to be encoded and transmitted to the decoder side besides the residual signal, where lots of coding bits are consumed. To further improve the performance of intra coding in Versatile Video Coding (VVC), an intelligent intra mode derivation method is proposed in this paper, termed as Deep Learning based Intra Mode Derivation (DLIMD). In specific, the process of intra mode derivation is formulated as a multi-class classification task, which aims to skip the module of intra mode signaling for coding bits reduction. The architecture of DLIMD is developed to adapt to different quantization parameter settings and variable coding blocks including non-square ones, where only one single trained model is required. Different from the existing deep learning based classification problems, the hand-crafted features are also fed into intra mode derivation network besides the learned features from feature learning network. To compete with traditional method, one additional binary flag is utilized in the video codec to indicate the selected scheme with RDO. Extensive experimental results reveal that the proposed method can achieve 2.28%, 1.74%, and 2.18% bit rate reduction on average for Y, U, and V components on the platform of VVC test model, which outperforms the state-of-the-art works.