Interpretable Visual Understanding with Cognitive Attention Network.

Interpretable Visual Understanding with Cognitive Attention Network.
复制标题

认知注意网络的可解释视觉理解。

DOI:
10.1007/978-3-030-86362-3_45
复制
发表时间:
2021-09
期刊:
Artificial neural networks, ICANN : international conference ... proceedings. International Conference on Artificial Neural Networks (European Neural Network Society)
影响因子:
--
通讯作者:
Ntoutsi E
Ntoutsi E
中科院分区:
其他
文献类型:
--
作者:
Tang X;Zhang W;Yu Y;Turner K;Derr T;Wang M;Ntoutsi E

文献摘要

相似文献

虽然图像理解在认知层面上取得了显著的进步,但可靠的视觉场景理解不仅需要在认知层面上对图像进行全面的理解,还需要在认知层面上对图像进行全面的理解,这就需要利用多源信息,学习不同层次的理解和广泛的常识知识。在本文中,我们提出了一种新的认知注意力网络(CAN)的视觉常识推理,以实现可解释的视觉理解。具体来说,我们首先介绍了一个图像-文本融合模块,融合信息的图像和文本集体。其次,设计了一种新的推理模块,用于对图像、查询和响应之间的常识进行编码。在大规模视觉常识推理(VCR)基准数据集上的实验证明了该方法的有效性。该实现可在https://github.com/tanjatang/CAN上公开获取
While image understanding on recognition-level has achieved remarkable advancements, reliable visual scene understanding requires comprehensive image understanding on recognition-level but also cognition-level, which calls for exploiting the multi-source information as well as learning different levels of understanding and extensive commonsense knowledge. In this paper, we propose a novel Cognitive Attention Network (CAN) for visual commonsense reasoning to achieve interpretable visual understanding. Specifically, we first introduce an image-text fusion module to fuse information from images and text collectively. Second, a novel inference module is designed to encode commonsense among image, query and response. Extensive experiments on large-scale Visual Commonsense Reasoning (VCR) benchmark dataset demonstrate the effectiveness of our approach. The implementation is publicly available at https://github.com/tanjatang/CAN