BabyTalk: Understanding and Generating Simple Image Descriptions

BabyTalk: Understanding and Generating Simple Image Descriptions
复制标题

DOI:
10.1109/tpami.2012.162
复制
发表时间:
2013-12-01
影响因子:
23.6
通讯作者:
Berg, Tamara L.
Berg, Tamara L.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Kulkarni, Girish;Premraj, Visruth;Berg, Tamara L.

文献摘要

被引文献

相似文献

我们提出了一个系统,自动生成自然语言描述的图像。该系统由两部分组成。第一部分,内容规划,平滑基于计算机视觉的检测和识别算法的输出,从大量的视觉描述性文本中挖掘统计数据,以确定用于描述图像的最佳内容词。第二步,表面实现,根据预测的内容和自然语言的一般统计数据选择词来构建自然语言句子。我们提出了多种方法的表面实现步骤和评估每个使用自动措施的相似性,人类生成的参考描述。我们还收集了强制选择人类之间的评价,从建议的生成系统和描述竞争的方法。所提出的系统是非常有效的生产相关的句子的图像。它还生成了比以前的工作更真实的特定图像内容的描述。
We present a system to automatically generate natural language descriptions from images. This system consists of two parts. The first part, content planning, smooths the output of computer vision-based detection and recognition algorithms with statistics mined from large pools of visually descriptive text to determine the best content words to use to describe an image. The second step, surface realization, chooses words to construct natural language sentences based on the predicted content and general statistics from natural language. We present multiple approaches for the surface realization step and evaluate each using automatic measures of similarity to human generated reference descriptions. We also collect forced choice human evaluations between descriptions from the proposed generation system and descriptions from competing approaches. The proposed system is very effective at producing relevant sentences for images. It also generates descriptions that are notably more true to the specific image content than previous work.