Audio-Visual Interpretable and Controllable Video Captioning
Audio-Visual Interpretable and Controllable Video Captioning
复制标题
DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Yapeng Tian;Chenxiao Guan;Justin Goodman;Marc Moore;Chenliang Xu
中科院分区:
文献类型:
--
作者:
Yapeng Tian;Chenxiao Guan;Justin Goodman;Marc Moore;Chenliang Xu
Video captioning [8] aims to automatically generate a natural language sentence to describe the dynamic, potentially complex multimodal scene inside of a video. Most of the previous works [8, 10] focus on exploring better visionand-language modelings and put less emphasis on the multimodal aspect of video captioning, where audio often reveals important in-scene and out-of-scene information and contributes to the language generation in its unique ways. For example, it adds the difficulty to describe the singing event by watching the audio-mute video in Fig. 1. Although this is not new to the multimedia community, many works over there aim to optimize the video captioning metrics with an uninterpretable fusion strategy. The basic questions as to what extent different modalities (auditory and visual) contribute to a particular sentence, and furthermore, to individual words in a sentence remain underexplored. It is our belief that unfolding these questions is valuable to making fundamental progress on video captioning. At first glance, it is seemingly impossible to answer the above questions. The first challenge is that there is no annotation denoting the individual contributions of the auditory or visual modality made to texts in any of the existing video captioning datasets—such a process is difficult to quantify without breakthroughs in Neurophysiology and Psychophysics. Instead, we study these questions from a computational perspective, where we mine signals from audio and video and compete their associations to text. The second challenge lies on the computational framework. Recurrent neural networks (RNNs) are widely used as decoders for video captioning. Despite the success in modeling sequential dependencies, RNN decoder-based architectures have inherent limitations to perform modalityinterpretable video captioning. When generating a word, besides using current provided/attended features and previous words, these models always exploit hidden states of the RNN decoder. The latter contains memorized information from different modalities, which makes the models impossible to disentangle the contributions from individual modalities for predicting the words. In this paper, we aim to disentangle the interplay of the two modalities and make the first attempt to interpretable audio-visual video captioning. Concretely, we propose a novel multimodal convolutional neural networkHumans: (1) people singing and dancing. (2) a group of people singing and dancing.