Weakly-Supervised Audio-Visual Video Parsing Toward Unified Multisensory Perception

Weakly-Supervised Audio-Visual Video Parsing Toward Unified Multisensory Perception
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yapeng Tian;Dingzeyu Li;Chenliang Xu
Yapeng Tian;Dingzeyu Li;Chenliang Xu
中科院分区:
其他
文献类型:
--
作者:
Yapeng Tian;Dingzeyu Li;Chenliang Xu

文献摘要

相似文献

利用和学习听觉和视觉模态是一个新兴的研究课题。近年来,在学习表示[1,2,6],分离视觉指示的声音[3,14],空间定位可见声源[8,12]以及时间定位视听同步片段[12]方面取得了进展。然而,过去的方法通常假设音频和视频数据总是相关的,甚至在时间上对齐。在实践中,当我们分析视频场景时,许多视频都有可听到的声音,这些声音起源于FoV之外,没有留下视觉对应,但仍然有助于整体理解,例如屏幕外运行的汽车和叙述者。这样的例子是如此普遍,这导致我们一些基本的问题:什么视频事件是可听的,可见的,和“audivisible”,这些事件在视频中的位置和时间,以及我们如何有效地检测它们?为了回答上述问题,我们提出并试图解决一个基本问题:视听视频解析,识别事件类别绑定到感官模态,同时,找到这样的事件何时开始和结束的时间边界(见图1)。然而,学习一个完全监督的视听视频解析模型需要密集注释的事件模态和类别标签与相应的事件开始和偏移,这将使标记过程非常昂贵和耗时。为了避免繁琐的标记,我们探索了任务的弱监督学习,它只需要对视频事件的存在或不存在进行稀疏标记。弱标签更容易注释,并且可以从网络视频中大规模收集。我们将弱监督下的视听视频解析归结为多模态多实例学习(MMIL)问题,并提出了一个新的框架来解决这个问题。具体地说,我们使用了一个新的混合注意力网络(HAN)来同时利用单模态和跨模态的时间上下文。我们开发了一个细心的MMIL池化方法,用于自适应地从不同的时间范围和方式聚合有用的音频和视频内容。此外,我们发现了模态偏差和噪声标签问题,并分别使用个人指导的学习机制和标签平滑来缓解它们[7]。为了方便我们的调查,我们收集了一个Look,Listen,and Parse(LLP)数据集,其中包含来自25个事件类别的11,849个YouTube视频片段。我们用稀疏的视频级事件标签标记它们进行训练。为了评估,我们标记了一组精确的标签,包括事件模态,事件A:篮球5s 10s
Utilizing and learning from both auditory and visual modalities is an emerging research topic. Recent years have seen progress in learning representations [1, 2, 6], separating visually indicated sounds [3, 14], spatially localizing visible sound sources [8, 12], and temporally localizing audio-visual synchronized segments [12]. However, past approaches usually assume audio and visual data are always correlated or even temporally aligned. In practice, when we analyze the video scene, many videos have audible sounds, which originate outside of the FoV, leaving no visual correspondences, but still contribute to the overall understanding, such as out-of-screen running cars and a narrating person. Such examples are so ubiquitous, which leads us to some basic questions: what video events are audible, visible, and “audivisible,” where and when are these events inside of a video, and how can we effectively detect them? To answer the above questions, we pose and try to tackle a fundamental problem: audio-visual video parsing that recognizes event categories bind to sensory modalities, and meanwhile, finds temporal boundaries of when such an event starts and ends (see Fig. 1). However, learning a fully supervised audio-visual video parsing model requires densely annotated event modality and category labels with corresponding event onsets and offsets, which will make the labeling process extremely expensive and time-consuming. To avoid tedious labeling, we explore weakly-supervised learning for the task, which only requires sparse labeling on the presence or absence of video events. The weak labels are easier to annotate and can be gathered in a large scale from web videos. We formulate the weakly-supervised audio-visual video parsing as a Multimodal Multiple Instance Learning (MMIL) problem and propose a new framework to solve it. Concretely, we use a new hybrid attention network (HAN) for leveraging unimodal and cross-modal temporal contexts simultaneously. We develop an attentive MMIL pooling method for adaptively aggregating useful audio and visual content from different temporal extent and modalities. Furthermore, we discover modality bias and noisy label issues and alleviate them with an individual-guided learning mechanism and label smoothing [7], respectively. To facilitate our investigations, we collect a Look, listen, and Parse (LLP) dataset that has 11, 849 YouTube video clips from 25 event categories. We label them with sparse video-level event labels for training. For evaluation, we label a set of precise labels, including event modalities, event A: Basketball 5s 10s