PSNet: Parallel Symmetric Network for Video Salient Object Detection

PSNet: Parallel Symmetric Network for Video Salient Object Detection
复制标题

DOI:
10.1109/tetci.2022.3220250
复制
发表时间:
2022-10
影响因子:
5.3
通讯作者:
Runmin Cong;Weiyu Song;Jianjun Lei;Guanghui Yue;Yao Zhao;Sam Kwong
Runmin Cong;Weiyu Song;Jianjun Lei;Guanghui Yue;Yao Zhao;Sam Kwong
中科院分区:
计算机科学2区
文献类型:
--
作者:
Runmin Cong;Weiyu Song;Jianjun Lei;Guanghui Yue;Yao Zhao;Sam Kwong

文献摘要

相似文献

对于视频显著对象检测(VSOD)任务,如何从外观通道和运动通道中挖掘信息一直是人们非常关注的话题。包括RGB外观流和光流运动流的双流结构被广泛用作VSOD任务的典型流水线,但现有方法通常只使用运动特征来单向引导外观特征或自适应但盲目地融合两种形态特征。然而,由于不全面和不具体的学习方案,这些方法在不同的场景中表现不佳。本文遵循更安全的建模思想,更全面地研究了外观通道和运动通道的重要性,提出了一种上下平行对称的VSOD网络PSNet。在聚集扩散增强(GDR)模块和跨通道细化与补码(CRC)模块的配合下,设置了两个具有不同主导通道的并行分支来实现完整的视频显著解码。最后,根据两个平行分支在不同场景中的不同重要程度,使用重要性感知融合(IPF)模块对其进行融合。在四个数据集基准测试上的实验表明,我们的方法取得了理想的和有竞争力的性能。
For the video salient object detection (VSOD) task, how to excavate the information from the appearance modality and the motion modality has always been a topic of great concern. The two-stream structure, including an RGB appearance stream and an optical flow motion stream, has been widely used as a typical pipeline for VSOD tasks, but the existing methods usually only use motion features to unidirectionally guide appearance features or adaptively but blindly fuse two modality features. However, these methods underperform in diverse scenarios due to the uncomprehensive and unspecific learning schemes. In this paper, following a more secure modeling philosophy, we deeply investigate the importance of appearance modality and motion modality in a more comprehensive way and propose a VSOD network with up and down parallel symmetry, named PSNet. Two parallel branches with different dominant modalities are set to achieve complete video saliency decoding with the cooperation of the Gather Diffusion Reinforcement (GDR) module and Cross-modality Refinement and Complement (CRC) module. Finally, we use the Importance Perception Fusion (IPF) module to fuse the features from two parallel branches according to their different importance in different scenarios. Experiments on four dataset benchmarks demonstrate that our method achieves desirable and competitive performance.