Learning to Ask Informative Sub-Questions for Visual Question Answering

Learning to Ask Informative Sub-Questions for Visual Question Answering
复制标题

DOI:
10.1109/cvprw56347.2022.00514
复制
发表时间:
2022-06
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
影响因子:
--
通讯作者:
Kohei Uehara;Nan Duan;Tatsuya Harada
Kohei Uehara;Nan Duan;Tatsuya Harada
中科院分区:
其他
文献类型:
--
作者:
Kohei Uehara;Nan Duan;Tatsuya Harada

文献摘要

相似文献

视觉问题推理(VQA)模型对于需要对世界知识进行推理的问题往往会做出错误的推理。最近的研究表明,训练VQA模型的问题,提供较低层次的感知信息沿着推理问题,提高性能。受此启发,我们提出了一种新的VQA模型,生成的问题,积极获得辅助感知信息,有助于正确的推理。我们的模型包括一个VQA模型回答问题,一个可视化的问题生成(VQG)模型生成的问题,和一个信息得分模型估计的信息量生成的问题包含,这是有用的回答原始问题。我们训练VQG模型,以最大限度地提高信息评分模型提供的“信息量”,以生成包含尽可能多的关于原始问题答案的信息的问题。我们的实验表明,通过将生成的问题及其答案作为附加信息输入到VQA模型中,它确实可以比基线模型更正确地预测答案。
VQA (Visual Question Answering) model tends to make incorrect inferences for questions that require reasoning over world knowledge. Recent study has shown that training VQA models with questions that provide lower-level perceptual information along with reasoning questions improves performance. Inspired by this, we propose a novel VQA model that generates questions to actively obtain auxiliary perceptual information useful for correct reasoning. Our model consists of a VQA model for answering questions, a Visual Question Generation (VQG) model for generating questions, and an Info-score model for estimating the amount of information the generated questions contain, which is useful in answering the original question. We train the VQG model to maximize the "informativeness" provided by the Info-score model to generate questions that contain as much information as possible, about the answer to the original question. Our experiments show that by inputting the generated questions and their answers as additional information to the VQA model, it can indeed predict the answer more correctly than the baseline model.