Interactive Re-ranking via Object Entropy-Guided Question Answering for Cross-Modal Image Retrieval

Interactive Re-ranking via Object Entropy-Guided Question Answering for Cross-Modal Image Retrieval
复制标题

DOI:
10.1145/3485042
复制
发表时间:
2022-03
期刊:
ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)
影响因子:
--
通讯作者:
Rintaro Yanagi;Ren Togo;Takahiro Ogawa;M. Haseyama
Rintaro Yanagi;Ren Togo;Takahiro Ogawa;M. Haseyama
中科院分区:
其他
文献类型:
--
作者:
Rintaro Yanagi;Ren Togo;Takahiro Ogawa;M. Haseyama

文献摘要

被引文献

相似文献

跨模态图像检索方法通过学习文本和图像之间的关系从查询文本中检索所需的图像。这种检索方法是实现查询准备简便性的最有效方法之一。当用户输入可用于唯一标识所需图像的查询文本时,最新的跨模式图像检索方法非常方便且准确。然而,现实中,用户经常输入模糊的查询文本,这些模糊的查询使得很难获得想要的图像。为了克服这些困难,在本研究中,我们提出了一种基于问答的新型交互式跨模态图像检索方法。该方法分析候选图像并向用户询问问题以获得可以缩小检索候选范围的信息。通过仅回答所提出的方法生成的问题,即使使用不明确的查询文本,用户也可以获得他们想要的图像。实验结果表明了该方法的有效性。
Cross-modal image-retrieval methods retrieve desired images from a query text by learning relationships between texts and images. Such a retrieval approach is one of the most effective ways of achieving the easiness of query preparation. Recent cross-modal image-retrieval methods are convenient and accurate when users input a query text that can be used to uniquely identify the desired image. However, in reality, users frequently input ambiguous query texts, and these ambiguous queries make it difficult to obtain desired images. To overcome these difficulties, in this study, we propose a novel interactive cross-modal image-retrieval method based on question answering. The proposed method analyzes candidate images and asks users questions to obtain information that can narrow down retrieval candidates. By only answering questions generated by the proposed method, users can reach their desired images, even when using an ambiguous query text. Experimental results show the proposed method’s effectiveness.