Investigating Robustness and Interpretability of Link Prediction via Adversarial Modifications

Investigating Robustness and Interpretability of Link Prediction via Adversarial Modifications
复制标题

DOI:
10.18653/v1/n19-1337
复制
发表时间:
2018-11
期刊:
--
影响因子:
--
通讯作者:
Pouya Pezeshkpour;Yifan Tian;Sameer Singh
Pouya Pezeshkpour;Yifan Tian;Sameer Singh
中科院分区:
其他
文献类型:
--
作者:
Pouya Pezeshkpour;Yifan Tian;Sameer Singh

文献摘要

相似文献

在嵌入空间中表示实体和关系是关系数据上机器学习的一种经过充分研究的方法。然而,现有的方法,主要集中在提高准确性和忽视其他方面,如鲁棒性和可解释性。在本文中,我们提出了对链接预测模型的对抗性修改:识别要添加到知识图中或从知识图中删除的事实,这些事实在模型重新训练后改变了对目标事实的预测。使用图的这些单个修改,我们确定了预测链接的最有影响力的事实,并评估了模型对添加假事实的敏感性。我们引入了一种有效的方法来估计这种修改的效果,当知识图发生变化时,通过近似嵌入的变化。为了避免对所有可能的事实进行组合搜索,我们训练了一个网络来解码嵌入到其相应的图组件,允许使用基于梯度的优化来识别对抗性修改。我们使用这些技术来评估链接预测模型的鲁棒性(通过测量对其他事实的敏感性),通过对预测最负责任的事实(通过识别最有影响力的邻居)来研究可解释性,并检测知识库中的不正确事实。
Representing entities and relations in an embedding space is a well-studied approach for machine learning on relational data. Existing approaches, however, primarily focus on improving accuracy and overlook other aspects such as robustness and interpretability. In this paper, we propose adversarial modifications for link prediction models: identifying the fact to add into or remove from the knowledge graph that changes the prediction for a target fact after the model is retrained. Using these single modifications of the graph, we identify the most influential fact for a predicted link and evaluate the sensitivity of the model to the addition of fake facts. We introduce an efficient approach to estimate the effect of such modifications by approximating the change in the embeddings when the knowledge graph changes. To avoid the combinatorial search over all possible facts, we train a network to decode embeddings to their corresponding graph components, allowing the use of gradient-based optimization to identify the adversarial modification. We use these techniques to evaluate the robustness of link prediction models (by measuring sensitivity to additional facts), study interpretability through the facts most responsible for predictions (by identifying the most influential neighbors), and detect incorrect facts in the knowledge base.