Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics

Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
复制标题

DOI:
10.1613/jair.3994
复制
发表时间:
2013-01-01
影响因子:
5
通讯作者:
Hockenmaier, Julia
Hockenmaier, Julia
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hodosh, Micah;Young, Peter;Hockenmaier, Julia

文献摘要

被引文献

相似文献

将图像与描述图像所描述内容的自然语言句子相关联的能力是图像理解的标志,也是基于句子的图像搜索等应用程序的先决条件。与图像搜索类似,我们建议将基于句子的图像标注框架为对给定的字幕池进行排序的任务。我们为基于句子的图像描述和搜索引入了一个新的基准集合,该集合由8000张图像组成,每张图像都配有5个不同的标题,这些标题提供了对显著实体和事件的清晰描述。我们介绍了一些在这项任务上表现相当好的系统,尽管它们只是基于可以在最小监督下获得的特征。我们的结果清楚地表明了对每个图像进行多个标题训练的重要性,以及捕获这些标题的句法(基于词序)和语义特征的重要性。我们还对该任务的人工和自动评估指标进行了深入的比较,并提出了低成本和大规模收集人工判断的策略,使我们能够通过附加的相关性判断来增强我们的收集,即哪些标题描述了哪些图像。我们的分析表明,考虑每个查询图像或句子的排序结果列表的指标明显比基于每个查询的单个响应的指标更健壮。此外,我们的研究表明,基于排名的图像描述系统的评估可能完全自动化。
The ability to associate images with natural language sentences that describe what is depicted in them is a hallmark of image understanding, and a prerequisite for applications such as sentence-based image search. In analogy to image search, we propose to frame sentence-based image annotation as the task of ranking a given pool of captions. We introduce a new benchmark collection for sentence-based image description and search, consisting of 8,000 images that are each paired with five different captions which provide clear descriptions of the salient entities and events. We introduce a number of systems that perform quite well on this task, even though they are only based on features that can be obtained with minimal supervision. Our results clearly indicate the importance of training on multiple captions per image, and of capturing syntactic (word order-based) and semantic features of these captions. We also perform an in-depth comparison of human and automatic evaluation metrics for this task, and propose strategies for collecting human judgments cheaply and on a very large scale, allowing us to augment our collection with additional relevance judgments of which captions describe which image. Our analysis shows that metrics that consider the ranked list of results for each query image or sentence are significantly more robust than metrics that are based on a single response per query. Moreover, our study suggests that the evaluation of ranking-based image description systems may be fully automated.