Query and Attention Augmentation for Knowledge-Based Explainable Reasoning

Query and Attention Augmentation for Knowledge-Based Explainable Reasoning
复制标题

DOI:
10.1109/cvpr52688.2022.01513
复制
发表时间:
2022-06
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yifeng Zhang;Ming Jiang;Qi Zhao
Yifeng Zhang;Ming Jiang;Qi Zhao
中科院分区:
其他
文献类型:
--
作者:
Yifeng Zhang;Ming Jiang;Qi Zhao

文献摘要

相似文献

可解释的视觉问答(VQA)模型已经开发了神经模块和基于查询的知识整合,以回答需要知识的问题。然而,大多数推理方法不能有效地生成查询或在推理过程中结合外部知识,这可能导致次优的结果。为了弥合这一研究差距,我们提出了查询和注意力增强,这是一种增强神经模块网络以联合推理视觉和外部知识的通用方法。为了在推理过程中考虑到这两个知识源,它将输入问题解析为一个函数程序,通过一种新的强化学习方法增强查询,并根据中间推理结果将增强的注意力共同引导到视觉和外部知识上。通过对多个VQA数据集的广泛实验,我们的方法在回答需要不同程度知识的问题时,表现出了比最先进的模型更高的性能、可解释性和可推广性。我们的源代码可在https://github.com/SuperJohnZhang/QAA上获得。
Explainable visual question answering (VQA) models have been developed with neural modules and query-based knowledge incorporation to answer knowledge-requiring questions. Yet, most reasoning methods cannot effectively generate queries or incorporate external knowledge during the reasoning process, which may lead to suboptimal results. To bridge this research gap, we present Query and Attention Augmentation, a general approach that augments neural module networks to jointly reason about visual and external knowledge. To take both knowledge sources into account during reasoning, it parses the input question into a functional program with queries augmented through a novel reinforcement learning method, and jointly directs augmented attention to visual and external knowledge based on intermediate reasoning results. With extensive experiments on multiple VQA datasets, our method demonstrates significant performance, explainability, and generalizability over state-of-the-art models in answering questions requiring different extents of knowledge. Our source code is available at https://github.com/SuperJohnZhang/QAA.