REX: Reasoning-aware and Grounded Explanation

REX: Reasoning-aware and Grounded Explanation
复制标题

DOI:
10.1109/cvpr52688.2022.01514
复制
发表时间:
2022-03
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Shi Chen;Qi Zhao
Shi Chen;Qi Zhao
中科院分区:
其他
文献类型:
--
作者:
Shi Chen;Qi Zhao

文献摘要

被引文献

相似文献

有效性和可解释性是可信人工智能系统的两个基本属性。视觉推理的最新研究大多致力于提高预测答案的准确性,而较少关注解释决策背后的基本原理。因此,他们通常利用虚假的偏差,而不是对视觉文本数据进行实际推理,并且尚未发展出通过考虑来自两种模式的关键信息来解释其决策的能力。本文旨在从三个不同的角度缩小这一差距:首先,我们定义了一种新型的多模态解释,通过逐步遍历推理过程和在图像中建立关键词来解释决策。我们开发了一个功能程序来顺序执行不同的推理步骤,并构建一个包含1,040,830个多模态解释的新数据集。其次,我们确定了在解释决策的视觉和文本模式中紧密耦合重要组件的关键需求,并提出了一种新的解释生成方法,该方法明确地模拟了单词和感兴趣区域之间的成对对应关系。它在很大程度上提高了视觉基础能力,从而增强了可解释性和推理性能。最后,利用我们的新数据和方法,我们进行了广泛的分析,以研究我们的解释在不同设置下的有效性,包括多任务学习和迁移学习。我们的代码和数据可在https://github.com/szzexpoi/rex上获得。
Effectiveness and interpretability are two essential properties for trustworthy AI systems. Most recent studies in visual reasoning are dedicated to improving the accuracy of predicted answers, and less attention is paid to explaining the rationales behind the decisions. As a result, they commonly take advantage of spurious biases instead of actually reasoning on the visual-textual data, and have yet developed the capability to explain their decision making by considering key information from both modalities. This paper aims to close the gap from three distinct perspectives: first, we define a new type of multi-modal explanations that explain the decisions by progressively traversing the reasoning process and grounding keywords in the images. We develop a functional program to sequentially ex-ecute different reasoning steps and construct a new dataset with 1,040,830 multi-modal explanations. Second, we iden-tify the critical need to tightly couple important components across the visual and textual modalities for explaining the decisions, and propose a novel explanation generation method that explicitly models the pairwise correspon-dence between words and regions of interest. It improves the visual grounding capability by a considerable margin, resulting in enhanced interpretability and reasoning performance. Finally, with our new data and method, we perform extensive analyses to study the effectiveness of our explanation under different settings, including multi-task learning and transfer learning. Our code and data are available at https://github.com/szzexpoi/rex.