A Symmetric Dual Encoding Dense Retrieval Framework for Knowledge-Intensive Visual Question Answering

A Symmetric Dual Encoding Dense Retrieval Framework for Knowledge-Intensive Visual Question Answering
复制标题

DOI:
10.1145/3539618.3591629
复制
发表时间:
2023-04
期刊:
Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
Alireza Salemi;Juan Altmayer Pizzorno;Hamed Zamani
Alireza Salemi;Juan Altmayer Pizzorno;Hamed Zamani
中科院分区:
其他
文献类型:
--
作者:
Alireza Salemi;Juan Altmayer Pizzorno;Hamed Zamani

文献摘要

相似文献

知识密集型视觉问答(KI-VQA)指的是回答一幅不在图像中的图像的问题。本文提出了一种新的KI-VQA任务流水线,由检索器和读取器组成。首先,我们介绍了DEDR,这是一种对称的双重编码密集检索框架,使用单模式(文本)和多模式编码器将文档和查询编码到共享的嵌入空间中。我们介绍了一种迭代的知识蒸馏方法,它弥合了这两个编码器中表示空间之间的差距。对两个成熟的Ki-VQA数据集,即OK-VQA和FVQA的广泛评估表明,DEDR在OK-VQA和FVQA上的性能分别比最先进的基线高11.6%和30.9%。利用DEDR检索到的段落,我们进一步介绍了一种编解码器多模式融合译码模型MM-FID,用于为Ki-VQA任务生成文本答案。MM-FID分别对问题、图像和检索到的每个段落进行编码,并在其解码器中联合使用所有段落。与文献中的竞争基线相比,该方法在OK-VQA和FVQA上的问答准确率分别提高了5.5%和8.5%。
Knowledge-Intensive Visual Question Answering (KI-VQA) refers to answering a question about an image whose answer does not lie in the image. This paper presents a new pipeline for KI-VQA tasks, consisting of a retriever and a reader. First, we introduce DEDR, a symmetric dual encoding dense retrieval framework in which documents and queries are encoded into a shared embedding space using uni-modal (textual) and multi-modal encoders. We introduce an iterative knowledge distillation approach that bridges the gap between the representation spaces in these two encoders. Extensive evaluation on two well-established KI-VQA datasets, i.e., OK-VQA and FVQA, suggests that DEDR outperforms state-of-the-art baselines by 11.6% and 30.9% on OK-VQA and FVQA, respectively. Utilizing the passages retrieved by DEDR, we further introduce MM-FiD, an encoder-decoder multi-modal fusion-in-decoder model, for generating a textual answer for KI-VQA tasks. MM-FiD encodes the question, the image, and each retrieved passage separately and uses all passages jointly in its decoder. Compared to competitive baselines in the literature, this approach leads to 5.5% and 8.5% improvements in terms of question answering accuracy on OK-VQA and FVQA, respectively.