Audio-Visual Event Localization via Recursive Fusion by Joint Co-Attention

Audio-Visual Event Localization via Recursive Fusion by Joint Co-Attention
复制标题

DOI:
10.1109/wacv48630.2021.00406
复制
发表时间:
2020-08
期刊:
2021 IEEE Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Bin Duan;Hao Tang;Wei Wang;Ziliang Zong;Guowei Yang;Yan Yan-Yan
Bin Duan;Hao Tang;Wei Wang;Ziliang Zong;Guowei Yang;Yan Yan-Yan
中科院分区:
其他
文献类型:
--
作者:
Bin Duan;Hao Tang;Wei Wang;Ziliang Zong;Guowei Yang;Yan Yan-Yan

文献摘要

被引文献

相似文献

视听事件定位任务的主要挑战在于如何有效地融合来自多模态的信息。最近的研究表明,注意机制是有益的融合过程。在本文中,我们提出了一种新的联合注意力机制与多模态融合方法的视听事件定位。特别是,我们提出了一个简洁而有效的架构,有效地学习表示从多种形式的联合方式。最初,视觉特征与听觉特征相结合,然后变成联合表示。接下来,我们利用联合表示来分别处理视觉特征和听觉特征。在这种联合注意的帮助下,产生了新的视觉和听觉特征,从而两个特征可以享受彼此的相互改进的好处。值得注意的是,联合共同关注单元是递归的,这意味着它可以被执行多次以逐渐获得更好的联合表示。在公共AVE数据集上的大量实验表明,该方法比现有方法取得了更好的结果。
The major challenge in audio-visual event localization task lies in how to fuse information from multiple modalities effectively. Recent works have shown that the attention mechanism is beneficial to the fusion process. In this paper, we propose a novel joint attention mechanism with multi-modal fusion methods for audio-visual event localization. Particularly, we present a concise yet valid architecture that effectively learns representations from multiple modalities in a joint manner. Initially, visual features are combined with auditory features and then turned into joint representations. Next, we make use of the joint representations to attend to visual features and auditory features, respectively. With the help of this joint co-attention, new visual and auditory features are produced, and thus both features can enjoy the mutually improved benefits from each other. It is worth noting that the joint co-attention unit is recursive meaning that it can be performed multiple times for obtaining better joint representations progressively. Extensive experiments on the public AVE dataset have shown that the proposed method achieves significantly better results than the state-of-the-art methods.