Deep Co-Attention Network for Multi-View Subspace Learning

Deep Co-Attention Network for Multi-View Subspace Learning
复制标题

DOI:
10.1145/3442381.3449801
复制
发表时间:
2021-02
期刊:
Proceedings of the Web Conference 2021
影响因子:
--
通讯作者:
Lecheng Zheng;Y. Cheng;Hongxia Yang;Nan Cao;Jingrui He
Lecheng Zheng;Y. Cheng;Hongxia Yang;Nan Cao;Jingrui He
中科院分区:
其他
文献类型:
--
作者:
Lecheng Zheng;Y. Cheng;Hongxia Yang;Nan Cao;Jingrui He

文献摘要

相似文献

许多现实世界的应用涉及来自多种模态的数据,因此呈现出视图异质性。例如,社交媒体上的用户建模可能会同时利用底层社交网络的拓扑结构和用户帖子的内容;在医学领域,多个视图可能是从不同姿势拍摄的X射线图像。到目前为止,已经提出了各种技术来取得有希望的结果,例如基于典型相关分析的方法等。同时,决策者能够理解这些方法的预测结果至关重要。例如,给定一个基于患者不同姿势的X射线图像的模型所提供的诊断结果,医生需要知道模型为什么做出这样的预测。然而,最先进的技术通常存在无法利用每个视图的互补信息以及以可解释的方式解释预测的问题。为了解决这些问题,在本文中,我们提出了一种用于多视图子空间学习的深度协同注意力网络,其目的是在对抗环境中提取公共信息和互补信息,并通过协同注意力机制为终端用户提供预测背后的可靠解释。特别是,它使用一种新颖的交叉重建损失,并利用标签信息通过将分类器纳入我们的模型来指导潜在表示的构建。这提高了潜在表示的质量并加快了收敛速度。最后,我们开发了一种高效的迭代算法来找到最优的编码器和鉴别器,并在合成数据集和现实世界数据集上进行了广泛评估。我们还进行了一个案例研究,以展示所提出的方法如何对图像数据集上的预测进行可靠解释。
Many real-world applications involve data from multiple modalities and thus exhibit the view heterogeneity. For example, user modeling on social media might leverage both the topology of the underlying social network and the content of the users’ posts; in the medical domain, multiple views could be X-ray images taken at different poses. To date, various techniques have been proposed to achieve promising results, such as canonical correlation analysis based methods, etc. In the meanwhile, it is critical for decision-makers to be able to understand the prediction results from these methods. For example, given the diagnostic result that a model provided based on the X-ray images of a patient at different poses, the doctor needs to know why the model made such a prediction. However, state-of-the-art techniques usually suffer from the inability to utilize the complementary information of each view and to explain the predictions in an interpretable manner. To address these issues, in this paper, we propose a deep co-attention network for multi-view subspace learning, which aims to extract both the common information and the complementary information in an adversarial setting and provide robust interpretations behind the prediction to the end-users via the co-attention mechanism. In particular, it uses a novel cross reconstruction loss and leverages the label information to guide the construction of the latent representation by incorporating the classifier into our model. This improves the quality of latent representation and accelerates the convergence speed. Finally, we develop an efficient iterative algorithm to find the optimal encoders and discriminator, which are evaluated extensively on synthetic and real-world data sets. We also conduct a case study to demonstrate how the proposed method robustly interprets the predictions on an image data set.