Audio-Visual Interpretable and Controllable Video Captioning

Audio-Visual Interpretable and Controllable Video Captioning
复制标题

DOI:
--
复制
发表时间:
2019
期刊:
--
影响因子:
--
通讯作者:
Yapeng Tian;Chenxiao Guan;Justin Goodman;Marc Moore;Chenliang Xu
Yapeng Tian;Chenxiao Guan;Justin Goodman;Marc Moore;Chenliang Xu
中科院分区:
其他
文献类型:
--
作者:
Yapeng Tian;Chenxiao Guan;Justin Goodman;Marc Moore;Chenliang Xu

文献摘要

被引文献

相似文献

视频字幕[8]旨在自动生成自然语言句子来描述视频中动态的、潜在复杂的多模态场景。之前的大多数作品 [8, 10] 都专注于探索更好的视觉和语言模型,而不太重视视频字幕的多模态方面,其中音频经常揭示重要的场景内和场景外信息,并以其独特的方式为语言生成做出贡献。例如,它增加了通过观看图1中的音频静音视频来描述歌唱事件的难度。虽然这对多媒体社区来说并不新鲜,但那里的许多工作旨在通过不可解释的融合策略来优化视频字幕指标。关于不同模态(听觉和视觉)对特定句子以及句子中各个单词的贡献程度的基本问题仍未得到充分探索。我们相信,解决这些问题对于在视频字幕方面取得根本性进展很有价值。乍一看,上述问题似乎无法回答。第一个挑战是,没有注释表明任何现有视频字幕数据集中的听觉或视觉模态对文本的单独贡献——如果神经生理学和心理物理学没有突破,这样的过程很难量化。相反,我们从计算的角度研究这些问题,从音频和视频中挖掘信号并比较它们与文本的关联。第二个挑战在于计算框架。循环神经网络 (RNN) 被广泛用作视频字幕的解码器。尽管在序列依赖性建模方面取得了成功,但基于 RNN 解码器的架构在执行模态可解释的视频字幕方面存在固有的局限性。在生成单词时,除了使用当前提供/关注的特征和先前的单词之外,这些模型总是利用 RNN 解码器的隐藏状态。后者包含来自不同模态的记忆信息,这使得模型无法区分单个模态对预测单词的贡献。在本文中,我们的目标是理清这两种模式的相互作用,并首次尝试可解释的视听视频字幕。具体来说,我们提出了一种新颖的多模态卷积神经网络Humans:(1)人们唱歌跳舞。 (2)一群人载歌载舞。
Video captioning [8] aims to automatically generate a natural language sentence to describe the dynamic, potentially complex multimodal scene inside of a video. Most of the previous works [8, 10] focus on exploring better visionand-language modelings and put less emphasis on the multimodal aspect of video captioning, where audio often reveals important in-scene and out-of-scene information and contributes to the language generation in its unique ways. For example, it adds the difficulty to describe the singing event by watching the audio-mute video in Fig. 1. Although this is not new to the multimedia community, many works over there aim to optimize the video captioning metrics with an uninterpretable fusion strategy. The basic questions as to what extent different modalities (auditory and visual) contribute to a particular sentence, and furthermore, to individual words in a sentence remain underexplored. It is our belief that unfolding these questions is valuable to making fundamental progress on video captioning. At first glance, it is seemingly impossible to answer the above questions. The first challenge is that there is no annotation denoting the individual contributions of the auditory or visual modality made to texts in any of the existing video captioning datasets—such a process is difficult to quantify without breakthroughs in Neurophysiology and Psychophysics. Instead, we study these questions from a computational perspective, where we mine signals from audio and video and compete their associations to text. The second challenge lies on the computational framework. Recurrent neural networks (RNNs) are widely used as decoders for video captioning. Despite the success in modeling sequential dependencies, RNN decoder-based architectures have inherent limitations to perform modalityinterpretable video captioning. When generating a word, besides using current provided/attended features and previous words, these models always exploit hidden states of the RNN decoder. The latter contains memorized information from different modalities, which makes the models impossible to disentangle the contributions from individual modalities for predicting the words. In this paper, we aim to disentangle the interplay of the two modalities and make the first attempt to interpretable audio-visual video captioning. Concretely, we propose a novel multimodal convolutional neural networkHumans: (1) people singing and dancing. (2) a group of people singing and dancing.