Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
复制标题

DOI:
10.1007/s11263-016-0981-7
复制
发表时间:
2017-05-01
影响因子:
19.5
通讯作者:
Li Fei-Fei
Li Fei-Fei
中科院分区:
计算机科学2区
文献类型:
--
作者:
Krishna, Ranjay;Zhu, Yuke;Li Fei-Fei

文献摘要

被引文献

相似文献

尽管在图像分类等感知任务方面取得了进展,但计算机在图像描述和问答等认知任务上的表现仍然不佳。认知是涉及不仅要识别,还要对我们的视觉世界进行推理的任务的核心。然而,用于处理图像中丰富内容以完成认知任务的模型仍然使用为感知任务设计的相同数据集进行训练。为了在认知任务上取得成功,模型需要理解图像中对象之间的相互作用和关系。当被问到“这个人骑的是什么交通工具?”时,计算机需要识别图像中的对象以及“骑(人,马车)”和“拉(马,马车)”的关系,才能正确回答“这个人乘坐的是一辆马拉的马车”。在本文中,我们提出了视觉基因组数据集,以便对这种关系进行建模。我们收集每个图像内对象、属性和关系的密集注释来学习这些模型。具体来说,我们的数据集包含超过10.8万张图像,其中每张图像平均有对象、属性以及对象之间的成对关系。我们将区域描述和问答对中的对象、属性、关系和名词短语规范为WordNet同义词集。这些注释共同构成了图像描述、对象、属性、关系和问答对最密集、最大的数据集。
Despite progress in perceptual tasks such as image classification, computers still perform poorly on cognitive tasks such as image description and question answering. Cognition is core to tasks that involve not just recognizing, but reasoning about our visual world. However, models used to tackle the rich content in images for cognitive tasks are still being trained using the same datasets designed for perceptual tasks. To achieve success at cognitive tasks, models need to understand the interactions and relationships between objects in an image. When asked "What vehicle is the person riding?", computers will need to identify the objects in an image as well as the relationships riding(man, carriage) and pulling(horse, carriage) to answer correctly that "the person is riding a horse-drawn carriage." In this paper, we present the Visual Genome dataset to enable the modeling of such relationships. We collect dense annotations of objects, attributes, and relationships within each image to learn these models. Specifically, our dataset contains over 108K images where each image has an average of objects, attributes, and pairwise relationships between objects. We canonicalize the objects, attributes, relationships, and noun phrases in region descriptions and questions answer pairs to WordNet synsets. Together, these annotations represent the densest and largest dataset of image descriptions, objects, attributes, relationships, and question answer pairs.