Multimodal Deep Neural Network with Image Sequence Features for Video Captioning

Multimodal Deep Neural Network with Image Sequence Features for Video Captioning
复制标题

DOI:
10.1109/ijcnn.2018.8489668
复制
发表时间:
2018-07
期刊:
2018 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Soichiro Oura;Tetsu Matsukawa;Einoshin Suzuki
Soichiro Oura;Tetsu Matsukawa;Einoshin Suzuki
中科院分区:
其他
文献类型:
--
作者:
Soichiro Oura;Tetsu Matsukawa;Einoshin Suzuki

文献摘要

相似文献

在本文中,我们提出了MDNNiSF(具有图像序列特征的多模态深度神经网络),用于生成给定视频片段的句子描述。最近提出的模型S2VT使用两个LSTM的堆栈来解决这个问题,并证明了高METEOR。然而,实验表明,S2VT有时会产生不准确的句子,这是很自然的,因为学习视觉和文本内容之间的关系具有挑战性。一个可能的原因是,视频字幕数据仍然很小的目的。我们试图通过将S2VT与NeuralTalk2集成来规避这一缺陷,NeuralTalk2用于图像字幕,并且由于其能够学习文本片段与图像片段之间的对齐,因此能够生成准确的描述。实验使用两个视频字幕数据,MSVD和MSRVTT,证明了我们的MDNNiSF在S2VT的有效性。例如,MDNNiSF在MSVD下达到了METEOR 0.344,比S2VT高21.5%。
In this paper, we propose MDNNiSF (Multimodal Deep Neural Network with image Sequence Features) for generating a sentence description of a given video clip. A recently proposed model, S2VT, uses a stack of two LSTMs to solve the problem and demonstrated high METEOR. However, experiments show that S2VT sometimes produces inaccurate sentences, which is quite natural due to the challenging nature of learning relationships between visual and textual contents. A possible reason is that the video caption data were still small for the purpose. We try to circumvent this flaw by integrating S2VT with NeuralTalk2, which is for image captioning and known to generate an accurate description due to its capability of learning alignments between text fragments to image fragments. Experiments using two video caption data, MSVD and MSRVTT, demonstrate the effectiveness of our MDNNiSF over S2VT. For example, MDNNiSF achieved METEOR 0.344, which is 21.5% higher than S2VT, with MSVD.