On the Robustness of Interpretability Methods

On the Robustness of Interpretability Methods
复制标题

DOI:
--
复制
发表时间:
2018-06
期刊:
ArXiv
影响因子:
--
通讯作者:
David Alvarez-Melis;T. Jaakkola
David Alvarez-Melis;T. Jaakkola
中科院分区:
其他
文献类型:
--
作者:
David Alvarez-Melis;T. Jaakkola

文献摘要

被引文献

相似文献

我们认为,解释的稳健性——即相似的输入应该产生相似的解释——是可解释性的关键需求。我们引入指标来量化鲁棒性,并证明当前方法根据这些指标表现不佳。最后,我们提出了在现有可解释性方法上增强鲁棒性的方法。
We argue that robustness of explanations---i.e., that similar inputs should give rise to similar explanations---is a key desideratum for interpretability. We introduce metrics to quantify robustness and demonstrate that current methods do not perform well according to these metrics. Finally, we propose ways that robustness can be enforced on existing interpretability approaches.