Discriminative Latent Semantic Graph for Video Captioning

Discriminative Latent Semantic Graph for Video Captioning
复制标题

DOI:
10.1145/3474085.3475519
复制
发表时间:
2021-08
期刊:
Proceedings of the 29th ACM International Conference on Multimedia
影响因子:
--
通讯作者:
Yang Bai;Junyan Wang;Yang Long;Bingzhang Hu;Yang Song;M. Pagnucco;Yu Guan
Yang Bai;Junyan Wang;Yang Long;Bingzhang Hu;Yang Song;M. Pagnucco;Yu Guan
中科院分区:
其他
文献类型:
--
作者:
Yang Bai;Junyan Wang;Yang Long;Bingzhang Hu;Yang Song;M. Pagnucco;Yu Guan

文献摘要

被引文献

相似文献

视频字幕旨在自动生成能够描述给定视频的视觉内容的自然语言句子。现有的生成模型,如编码器-解码器框架,不能显式地从复杂的时空数据中探索对象级交互和帧级信息,以生成语义丰富的字幕。我们的主要贡献是确定未来的视频摘要任务的联合框架中的三个关键问题。1)增强的对象建议:我们提出了一种新的条件图,可以融合时空信息到潜在的对象建议。2)视觉知识:潜在的建议聚合,提出了动态提取视觉词与更高的语义水平。3)句子验证:提出了一种新的判别式语言验证器来验证生成的字幕,以便有效地保留关键语义概念。我们在两个公共数据集(MVSD和MSR-VTT)上的实验表明,在所有指标上,特别是BLEU-4和CIDEr上,与最先进的方法相比,都有显著的改进。我们的代码可在https://github.com/baiyang4/D-LSG-Video-Caption上获得。
Video captioning aims to automatically generate natural language sentences that can describe the visual contents of a given video. Existing generative models like encoder-decoder frameworks cannot explicitly explore the object-level interactions and frame-level information from complex spatio-temporal data to generate semantic-rich captions. Our main contribution is to identify three key problems in a joint framework for future video summarization tasks. 1) Enhanced Object Proposal: we propose a novel Conditional Graph that can fuse spatio-temporal information into latent object proposal. 2) Visual Knowledge: Latent Proposal Aggregation is proposed to dynamically extract visual words with higher semantic levels. 3) Sentence Validation: A novel Discriminative Language Validator is proposed to verify generated captions so that key semantic concepts can be effectively preserved. Our experiments on two public datasets (MVSD and MSR-VTT) manifest significant improvements over state-of-the-art approaches on all metrics, especially for BLEU-4 and CIDEr. Our code is available at https://github.com/baiyang4/D-LSG-Video-Caption.