End-to-end Knowledge Retrieval with Multi-modal Queries

End-to-end Knowledge Retrieval with Multi-modal Queries
复制标题

DOI:
10.48550/arxiv.2306.00424
复制
发表时间:
2023-06
期刊:
--
影响因子:
--
通讯作者:
Man Luo;Zhiyuan Fang;Tejas Gokhale;Yezhou Yang;Chitta Baral
Man Luo;Zhiyuan Fang;Tejas Gokhale;Yezhou Yang;Chitta Baral
中科院分区:
其他
文献类型:
--
作者:
Man Luo;Zhiyuan Fang;Tejas Gokhale;Yezhou Yang;Chitta Baral

文献摘要

相似文献

我们研究知识检索与多模态查询,即查询包含信息分裂跨图像和文本输入,一个具有挑战性的任务,不同于以前的工作跨模态检索。我们策划了一个名为ReMuQ的新数据集,用于对这项任务的进展进行基准测试。ReMuQ需要一个系统,通过整合文本和图像查询的内容来从大型语料库中检索知识。我们引入了一个检索模型“ReViz”,它可以直接处理输入的文本和图像,以端到端的方式检索相关知识,而不依赖于中间模块,如对象检测器或字幕生成器。我们引入了一个新的预训练任务,该任务对于使用多模态查询学习知识检索是有效的,并且还提高了下游任务的性能。我们展示了在零拍摄设置下在两个数据集(ReMuQ和OK-VQA)上检索的上级性能,以及在这些数据集上进行微调时的进一步改进。
We investigate knowledge retrieval with multi-modal queries, i.e. queries containing information split across image and text inputs, a challenging task that differs from previous work on cross-modal retrieval. We curate a new dataset called ReMuQ for benchmarking progress on this task. ReMuQ requires a system to retrieve knowledge from a large corpus by integrating contents from both text and image queries. We introduce a retriever model “ReViz” that can directly process input text and images to retrieve relevant knowledge in an end-to-end fashion without being dependent on intermediate modules such as object detectors or caption generators. We introduce a new pretraining task that is effective for learning knowledge retrieval with multimodal queries and also improves performance on downstream tasks. We demonstrate superior performance in retrieval on two datasets (ReMuQ and OK-VQA) under zero-shot settings as well as further improvements when finetuned on these datasets.