Weakly-Supervised Audio-Visual Video Parsing Toward Unified Multisensory Perception
Weakly-Supervised Audio-Visual Video Parsing Toward Unified Multisensory Perception
复制标题
DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Yapeng Tian;Dingzeyu Li;Chenliang Xu
中科院分区:
文献类型:
--
作者:
Yapeng Tian;Dingzeyu Li;Chenliang Xu
Utilizing and learning from both auditory and visual modalities is an emerging research topic. Recent years have seen progress in learning representations [1, 2, 6], separating visually indicated sounds [3, 14], spatially localizing visible sound sources [8, 12], and temporally localizing audio-visual synchronized segments [12]. However, past approaches usually assume audio and visual data are always correlated or even temporally aligned. In practice, when we analyze the video scene, many videos have audible sounds, which originate outside of the FoV, leaving no visual correspondences, but still contribute to the overall understanding, such as out-of-screen running cars and a narrating person. Such examples are so ubiquitous, which leads us to some basic questions: what video events are audible, visible, and “audivisible,” where and when are these events inside of a video, and how can we effectively detect them? To answer the above questions, we pose and try to tackle a fundamental problem: audio-visual video parsing that recognizes event categories bind to sensory modalities, and meanwhile, finds temporal boundaries of when such an event starts and ends (see Fig. 1). However, learning a fully supervised audio-visual video parsing model requires densely annotated event modality and category labels with corresponding event onsets and offsets, which will make the labeling process extremely expensive and time-consuming. To avoid tedious labeling, we explore weakly-supervised learning for the task, which only requires sparse labeling on the presence or absence of video events. The weak labels are easier to annotate and can be gathered in a large scale from web videos. We formulate the weakly-supervised audio-visual video parsing as a Multimodal Multiple Instance Learning (MMIL) problem and propose a new framework to solve it. Concretely, we use a new hybrid attention network (HAN) for leveraging unimodal and cross-modal temporal contexts simultaneously. We develop an attentive MMIL pooling method for adaptively aggregating useful audio and visual content from different temporal extent and modalities. Furthermore, we discover modality bias and noisy label issues and alleviate them with an individual-guided learning mechanism and label smoothing [7], respectively. To facilitate our investigations, we collect a Look, listen, and Parse (LLP) dataset that has 11, 849 YouTube video clips from 25 event categories. We label them with sparse video-level event labels for training. For evaluation, we label a set of precise labels, including event modalities, event A: Basketball 5s 10s