Towards Robust Interpretability with Self-Explaining Neural Networks

Towards Robust Interpretability with Self-Explaining Neural Networks
复制标题

DOI:
--
复制
发表时间:
2018-06
期刊:
ArXiv
影响因子:
--
通讯作者:
David Alvarez-Melis;T. Jaakkola
David Alvarez-Melis;T. Jaakkola
中科院分区:
其他
文献类型:
--
作者:
David Alvarez-Melis;T. Jaakkola

文献摘要

被引文献

相似文献

最近关于复杂机器学习模型可解释性的工作主要集中于估计先前围绕特定预测训练的模型的后验解释。可解释性在学习过程中已经发挥关键作用的自解释模型受到的关注要少得多。我们提出了一般解释的三个必要条件——明确性、忠实性和稳定性——并表明现有方法不能满足这些要求。作为回应,我们分阶段设计自解释模型,逐步将线性分类器推广到复杂但架构明确的模型。通过专门针对此类模型定制的正则化来增强忠实性和稳定性。各种基准数据集的实验结果表明,我们的框架为协调模型复杂性和可解释性提供了一个有希望的方向。
Most recent work on interpretability of complex machine learning models has focused on estimating a posteriori explanations for previously trained models around specific predictions. Self-explaining models where interpretability plays a key role already during learning have received much less attention. We propose three desiderata for explanations in general – explicitness, faithfulness, and stability – and show that existing methods do not satisfy them. In response, we design self-explaining models in stages, progressively generalizing linear classifiers to complex yet architecturally explicit models. Faithfulness and stability are enforced via regularization specifically tailored to such models. Experimental results across various benchmark datasets show that our framework offers a promising direction for reconciling model complexity and interpretability.