Recallable Question Answering-Based Re-Ranking Considering Semantic Region for Cross-Modal Retrieval

Recallable Question Answering-Based Re-Ranking Considering Semantic Region for Cross-Modal Retrieval
复制标题

DOI:
10.1109/ojsp.2023.3238280
复制
发表时间:
2023
影响因子:
2.8
通讯作者:
Rintaro Yanagi;Ren Togo;Takahiro Ogawa;M. Haseyama
Rintaro Yanagi;Ren Togo;Takahiro Ogawa;M. Haseyama
中科院分区:
--
文献类型:
--
作者:
Rintaro Yanagi;Ren Togo;Takahiro Ogawa;M. Haseyama

文献摘要

相似文献

最近提出了基于问答(QA)的跨模态检索重新排序方法,以进一步缩小相似的候选图像。传统的基于QA的重排序方法通过分析候选图像向用户提供问题,并根据用户的反馈对初始检索结果进行重排序。与这些发展相反,仅仅关注性能改进使得很难有效地引出用户的检索意图。为了实现更有用的基于质量保证的重排序,需要考虑用户交互以引出用户的检索意图。在本文中,我们提出了一个基于QA的重新排序方法,考虑两个重要因素,以引起用户的检索意图:查询图像的相关性和召回。考虑查询图像相关性使得能够仅关注与所提供的查询文本相关的候选图像,而关注可回忆性使得用户能够容易地回答所提供的问题。通过这些步骤,我们的方法可以有效地和有效地引发用户的检索意图。使用Microsoft Common Objects in Context和计算构造的包含相似候选图像的数据集的实验结果表明,该方法可以提高跨模态检索方法和基于QA的重排序方法的性能。
Question answering (QA)-based re-ranking methods for cross-modal retrieval have been recently proposed to further narrow down similar candidate images. The conventional QA-based re-ranking methods provide questions to users by analyzing candidate images, and the initial retrieval results are re-ranked based on the user's feedback. Contrary to these developments, only focusing on performance improvement makes it difficult to efficiently elicit the user's retrieval intention. To realize more useful QA-based re-ranking, considering the user interaction for eliciting the user's retrieval intention is required. In this paper, we propose a QA-based re-ranking method with considering two important factors for eliciting the user's retrieval intention: query-image relevance and recallability. Considering the query-image relevance enables to only focus on the candidate images related to the provided query text, while, focusing on the recallability enables users to easily answer the provided question. With these procedures, our method can efficiently and effectively elicit the user's retrieval intention. Experimental results using Microsoft Common Objects in Context and computationally constructed dataset including similar candidate images show that our method can improve the performance of the cross-modal retrieval methods and the QA-based re-ranking methods.