Self-learn to Explain Siamese Networks Robustly

Self-learn to Explain Siamese Networks Robustly
复制标题

DOI:
10.1109/icdm51629.2021.00116
复制
发表时间:
2021-09
期刊:
2021 IEEE International Conference on Data Mining (ICDM)
影响因子:
--
通讯作者:
Chao Chen;Yifan Shen;Guixiang Ma;Xiangnan Kong;S. Rangarajan;Xi Zhang;Sihong Xie
Chao Chen;Yifan Shen;Guixiang Ma;Xiangnan Kong;S. Rangarajan;Xi Zhang;Sihong Xie
中科院分区:
其他
文献类型:
--
作者:
Chao Chen;Yifan Shen;Guixiang Ma;Xiangnan Kong;S. Rangarajan;Xi Zhang;Sihong Xie

文献摘要

被引文献

相似文献

学习比较两个对象在应用中至关重要,特别是当标记数据稀缺且不平衡时。由于这些应用程序可能涉及人类并做出高风险决策,因此解释学习模型至关重要。我们的目标是研究广泛用于学习比较的连体网络(SN)的事后解释。与具有单个输入实例的架构相比,我们描述了由于 SN 中额外的比较对象而导致的基于梯度的解释的不稳定性。我们使用自学习基于未标记数据优化全局不变性,以促进个体输入的局部解释的稳定性。这种不变性导致了约束优化问题,可以使用梯度下降-上升 (GDA) 或由 SGD 解决的 KL 散度正则化无约束优化来解决。当目标函数由于孪生架构而非凸时,我们提供收敛证明。神经科学和化学工程的表格和图形数据的结果表明,我们的局部解释在优化解释的忠实性和简单性的同时,有力地尊重了自学的不变性。我们进一步通过实验证明了 GDA 的收敛性。
Learning to compare two objects are essential in applications, especially when labeled data are scarce and imbalanced. As these applications can involve humans and make high-stake decisions, it is critical to explain the learned models. We aim to study post-hoc explanations of Siamese networks (SN) widely used in learning to compare. We characterize the instability of gradient-based explanations due to the additional compared object in SN, in contrast to architectures with a single input instance. We optimize for global invariance based on unlabeled data using self-learning to promote the stability of local explanations for individual input. The invariance leads to constrained optimization problems that can be solved using gradient descent-ascent (GDA), or KL-divergence regularized unconstrained optimization solved by SGD. We provide convergence proofs when the objective functions are nonconvex due to the Siamese architecture. Results on tabular and graph data from neuroscience and chemical engineering show that our local explanations robustly respects the self-learned invariance while optimizing the explanation faithfulness and simplicity. We further demonstrate the convergence of GDA experimentally.