Natural Language Object Retrieval

Natural Language Object Retrieval
复制标题

DOI:
10.1109/cvpr.2016.493
复制
发表时间:
2015-11
期刊:
2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Ronghang Hu;Huazhe Xu;Marcus Rohrbach;Jiashi Feng;Kate Saenko;Trevor Darrell
Ronghang Hu;Huazhe Xu;Marcus Rohrbach;Jiashi Feng;Kate Saenko;Trevor Darrell
中科院分区:
其他
文献类型:
--
作者:
Ronghang Hu;Huazhe Xu;Marcus Rohrbach;Jiashi Feng;Kate Saenko;Trevor Darrell

文献摘要

被引文献

相似文献

在本文中,我们讨论了自然语言对象检索的任务,即基于对象的自然语言查询在给定的图像中定位目标对象。自然语言对象检索不同于基于文本的图像检索任务,因为它涉及场景内对象的空间信息和全局场景上下文。为了解决这一问题,我们提出了一种新的空间上下文递归ConvNet(SCRC)模型作为对象检索候选盒的评分函数,将空间配置和全局场景级上下文信息整合到网络中。该模型通过递归网络处理查询文本、局部图像描述符、空间配置和全局上下文特征,输出查询文本在每个候选框上的概率作为该框的得分,并将视觉语言知识从图像字幕域转移到我们的任务中。实验结果表明,我们的方法有效地利用了局部和全局信息,在不同的数据集和场景上显著优于以往的基线方法,并且可以利用大规模的视觉和语言数据集来进行知识转移。
In this paper, we address the task of natural language object retrieval, to localize a target object within a given image based on a natural language query of the object. Natural language object retrieval differs from text-based image retrieval task as it involves spatial information about objects within the scene and global scene context. To address this issue, we propose a novel Spatial Context Recurrent ConvNet (SCRC) model as scoring function on candidate boxes for object retrieval, integrating spatial configurations and global scene-level contextual information into the network. Our model processes query text, local image descriptors, spatial configurations and global context features through a recurrent network, outputs the probability of the query text conditioned on each candidate box as a score for the box, and can transfer visual-linguistic knowledge from image captioning domain to our task. Experimental results demonstrate that our method effectively utilizes both local and global information, outperforming previous baseline methods significantly on different datasets and scenarios, and can exploit large scale vision and language datasets for knowledge transfer.