A Corpus for Reasoning about Natural Language Grounded in Photographs

A Corpus for Reasoning about Natural Language Grounded in Photographs
复制标题

DOI:
10.18653/v1/p19-1644
复制
发表时间:
2018-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Alane Suhr;Stephanie Zhou;Iris Zhang;Huajun Bai;Yoav Artzi
Alane Suhr;Stephanie Zhou;Iris Zhang;Huajun Bai;Yoav Artzi
中科院分区:
其他
文献类型:
--
作者:
Alane Suhr;Stephanie Zhou;Iris Zhang;Huajun Bai;Yoav Artzi

文献摘要

相似文献

我们介绍了一个新的数据集,用于自然语言和图像的联合推理,重点是语义多样性、组合性和视觉推理挑战。这些数据包含107,292个英语句子的例子和网络照片。这项任务是确定关于一对照片的自然语言说明是否属实。我们使用一组视觉丰富的图像和一个比较对比的任务来众包数据,以引出不同的语言。定性分析表明,数据需要成分联合推理,包括关于数量、比较和关系的推理。使用最先进的视觉推理方法进行评估显示,数据提出了强大的挑战。
We introduce a new dataset for joint reasoning about natural language and images, with a focus on semantic diversity, compositionality, and visual reasoning challenges. The data contains 107,292 examples of English sentences paired with web photographs. The task is to determine whether a natural language caption is true about a pair of photographs. We crowdsource the data using sets of visually rich images and a compare-and-contrast task to elicit linguistically diverse language. Qualitative analysis shows the data requires compositional joint reasoning, including about quantities, comparisons, and relations. Evaluation using state-of-the-art visual reasoning methods shows the data presents a strong challenge.