Telling Left From Right: Learning Spatial Correspondence of Sight and Sound

Telling Left From Right: Learning Spatial Correspondence of Sight and Sound
复制标题

区分左右:学习视觉和声音的空间对应

DOI:
10.1109/cvpr42600.2020.00995
复制
发表时间:
2020
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
J. Salamon
J. Salamon
中科院分区:
--
文献类型:
--
作者:
Karren D. Yang;Bryan C. Russell;J. Salamon

文献摘要

被引文献

相似文献

自我监督视听学习旨在通过利用视觉和音频输入之间的对应关系来捕获有用的视频表示。现有的方法主要集中在感官流之间的语义信息的匹配。我们提出了一种新的自我监督的任务,利用正交原则:匹配音频流中的空间信息的声音源在视觉流的位置。我们的方法简单而有效。我们训练一个模型来确定左右音频通道是否被翻转,迫使它在视觉和音频流中进行空间定位。为了训练和评估我们的方法,我们引入了一个大规模的视频数据集YouTube-ASMR-300 K,其中空间音频包含超过900小时的镜头。我们证明,理解空间对应性使模型能够在三个视听任务上表现得更好,在不利用空间音频线索的监督和自我监督基线上实现定量增益。我们还展示了如何将我们的自我监督方法扩展到具有ambisonic音频的360度视频。
Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic information between the sensory streams. We propose a novel self-supervised task to leverage an orthogonal principle: matching spatial information in the audio stream to the positions of sound sources in the visual stream. Our approach is simple yet effective. We train a model to determine whether the left and right audio channels have been flipped, forcing it to reason about spatial localization across the visual and audio streams. To train and evaluate our method, we introduce a large-scale video dataset, YouTube-ASMR-300K, with spatial audio comprising over 900 hours of footage. We demonstrate that understanding spatial correspondence enables models to perform better on three audio-visual tasks, achieving quantitative gains over supervised and self-supervised baselines that do not leverage spatial audio cues. We also show how to extend our self-supervised approach to 360 degree videos with ambisonic audio.