Audio-visual aligned saliency model for omnidirectional video with implicit neural representation learning

Audio-visual aligned saliency model for omnidirectional video with implicit neural representation learning
复制标题

DOI:
10.1007/s10489-023-04714-1
复制
发表时间:
2023-06
影响因子:
5.3
通讯作者:
Dandan Zhu;X. Shao;Kaiwei Zhang;Xiongkuo Min;Guangtao Zhai;Xiaokang Yang
Dandan Zhu;X. Shao;Kaiwei Zhang;Xiongkuo Min;Guangtao Zhai;Xiaokang Yang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Dandan Zhu;X. Shao;Kaiwei Zhang;Xiongkuo Min;Guangtao Zhai;Xiaokang Yang

文献摘要

相似文献

由于全方位视频(ODV)中的音频信息被充分挖掘和利用,现有的视听显著性模型的性能得到了显着和显着的改善。然而,这些模型仍处于起步阶段,在视觉和听觉模态之间建模人类注意力存在两个重要问题:(1)听觉和视觉模态之间的时间不对齐问题很少被考虑;(2)大多数视听显着性模型是音频内容属性不可知的。因此,他们需要学习具有精细细节的音频特征。本文提出了一种新的视听对齐显着性(AVAS)模型,可以同时解决上述两个问题,在一个有效的端到端的训练方式。为了解决两种模态之间的时间不对齐问题,在音频流上采用Hanning窗方法,以每单位时间(帧-时间间隔)截断音频信号,以匹配相应持续时间的视觉信息流,该方法可以捕获跨时间步长的两种模态的潜在相关性,从而促进视听特征融合。针对音频内容属性不可知的问题,提出了一种基于隐式神经表示(INR)的周期性音频编码方法,将音频采样点映射到相应的音频频率值,从而更好地区分和解释音频内容属性。在基准数据集上进行了全面的实验和详细的消融分析,以证明所提出的模型的有效性。实验结果表明,该模型始终优于其他竞争对手的大幅度。
Since the audio information is fully explored and leveraged in omnidirectional videos (ODVs), the performance of existing audio-visual saliency models has been improving dramatically and significantly. However, these models are still in their infancy stages, and there are two significant issues in modeling human attention between visual and auditory modalities: (1) Temporal non-alignment problem between auditory and visual modalities is rarely considered; (2) Most audio-visual saliency models are audio content attributes-agnostic. Thus, they need to learn audio features with fine details. This paper proposes a novel audio-visual aligned saliency (AVAS) model that can simultaneously tackle two issues as mentioned above in an effective end-to-end training manner. In order to solve the temporal non-alignment problem between the two modalities, a Hanning window method is employed on the audio stream to truncate the audio signal per unit time (frame-time interval) to match the visual information stream of the corresponding duration, which can capture the potential correlation of two modalities across time steps and facilitate audio-visual features fusion. Regarding the audio content attribute-agnostic issue, an effective periodic audio encoding method is proposed based on implicit neural representation (INR) to map audio sampling points to their corresponding audio frequency values, which can better discriminate and interpret audio content attributes. Comprehensive experiments and detailed ablation analyses are performed on the benchmark dataset to demonstrate the efficacy of the proposed model. The experimental results indicate that the proposed model consistently outperforms other competitors by a large margin.