Multi-Stage Spatial and Frequency Feature Fusion using Transformer in CNN-Based In-Loop Filter for VVC

Multi-Stage Spatial and Frequency Feature Fusion using Transformer in CNN-Based In-Loop Filter for VVC
复制标题

DOI:
10.1109/pcs56426.2022.10017998
复制
发表时间:
2022-12
期刊:
2022 Picture Coding Symposium (PCS)
影响因子:
--
通讯作者:
B. Kathariya;Zhu Li;Hongtao Wang;Mohammad Coban
B. Kathariya;Zhu Li;Hongtao Wang;Mohammad Coban
中科院分区:
其他
文献类型:
--
作者:
B. Kathariya;Zhu Li;Hongtao Wang;Mohammad Coban

文献摘要

相似文献

通用视频编码(VVC)/H.266是高效视频编码(HEVC)/H.255和高级视频编码(AVC)/H.264的视频编码后继者,具有显著的技术和编码改进。尽管如此,它遵循传统的基于块的混合视频编码方案类似于它的前辈。结果是,重构的图像包含压缩伪影。默认情况下,VVC具有环路滤波器来纠正畸形,但这些手工制作的滤波器提供次优性能。在这项工作中,我们设计了一种新的卷积神经网络(CNN)来取代VVC的内置环路滤波器。提出的基于卷积神经网络的环路滤波器利用多级谱Transformer(MST++)的改进的谱多级头自注意(S-MSA)层在多个阶段融合分别从像素及其离散余弦变换(DCT)应用输入中提取的空间和频率分解特征。我们将所提出的网络命名为MSTFNet,其中前三个字母代表MST++,F代表融合。由于多阶段的特征融合操作,所提出的CNN作为一个强大的学习环路滤波器,显着优于以前的方法。我们的实验结果表明,所提出的方法可以实现编码的平均Bjøntegaard增量(BD)比特率节省高达10.31%的亮度(Y)分量下的所有帧内(AI)配置。
Versatile Video Coding (VVC)/H.266 is a video coding successor to High Efficiency Video Coding (HEVC)/H.255 and Advanced Video Coding (AVC)/H.264 with significant technical and coding improvement. Nonetheless, it follows the conventional block-based hybrid video coding scheme similar to its predecessors. The consequence is, that the reconstructed picture contains compression artifacts. VVC, by default, has in-loop filters to correct the deformities but these handcrafted filters offer suboptimal performance. In this work, we designed a novel convolutional neural network (CNN) to replace the inbuilt in-loop filter of VVC. The proposed CNN-based in-loop filter utilizes a modified Spectral-wise Multi-Head Self-Attention (S-MSA) layer of Multistage Spectral-wise Transformer (MST++) at multiple stages to fuse spatial and frequency-decomposed features extracted from pixel and its discrete-cosine-transform (DCT) applied input respectively. We named the proposed network MSTFNet where the first three letters represent MST++ and F stands for fusion. Because of the multi-stage feature fusion operation, the proposed CNN acts as a powerful learned in-loop filter that significantly outperforms previous methods. Our experimental results show that the proposed method can achieve coding improvements up to 10.31% on average Bjøntegaard Delta (BD)-Bitrate savings under all-intra (AI) configurations for the luma (Y) component.