Integrating Language and Vision to Generate Natural Language Descriptions of Videos in the Wild

Integrating Language and Vision to Generate Natural Language Descriptions of Videos in the Wild
复制标题

DOI:
--
复制
发表时间:
2014-08
期刊:
--
影响因子:
--
通讯作者:
Jesse Thomason;Subhashini Venugopalan;S. Guadarrama;Kate Saenko;R. Mooney
Jesse Thomason;Subhashini Venugopalan;S. Guadarrama;Kate Saenko;R. Mooney
中科院分区:
其他
文献类型:
--
作者:
Jesse Thomason;Subhashini Venugopalan;S. Guadarrama;Kate Saenko;R. Mooney

文献摘要

被引文献

相似文献

本文集成了自然语言处理和计算机视觉技术,以提高对现实世界视频中实体和活动的识别和描述。我们提出了一种通过使用因子图将视觉检测与语言统计相结合来生成视频文本描述的策略。我们使用最先进的视觉识别系统来获得对视频中存在的实体、活动和场景的置信度。我们的因子图模型将这些检测置信度与从文本语料库中挖掘的概率知识相结合,以估计最可能的主语、动词、宾语和地点。 YouTube 视频上的结果表明,与单独使用视觉系统以及之前的 n-gram 语言建模方法相比,我们的方法改进了这些潜在的、多样化的句子成分的联合检测以及某些单独成分的检测。联合检测使我们能够自动生成更准确、更丰富的视频句子描述,并包含多种可能的内容。
This paper integrates techniques in natural language processing and computer vision to improve recognition and description of entities and activities in real-world videos. We propose a strategy for generating textual descriptions of videos by using a factor graph to combine visual detections with language statistics. We use state-of-the-art visual recognition systems to obtain confidences on entities, activities, and scenes present in the video. Our factor graph model combines these detection confidences with probabilistic knowledge mined from text corpora to estimate the most likely subject, verb, object, and place. Results on YouTube videos show that our approach improves both the joint detection of these latent, diverse sentence components and the detection of some individual components when compared to using the vision system alone, as well as over a previous n-gram language-modeling approach. The joint detection allows us to automatically generate more accurate, richer sentential descriptions of videos with a wide array of possible content.