Towards falsifiable interpretability research

Towards falsifiable interpretability research
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Matthew L. Leavitt;Ari S. Morcos
Matthew L. Leavitt;Ari S. Morcos
中科院分区:
其他
文献类型:
--
作者:
Matthew L. Leavitt;Ari S. Morcos

文献摘要

被引文献

相似文献

理解深度神经网络(dnn)的决策和机制的方法通常依赖于通过强调单个示例的感觉或语义特征来建立直觉。例如,方法旨在可视化输入的对网络决策“重要”的组成部分,或测量单个神经元的语义属性。在这里,我们认为可解释性研究遭受了对基于直觉的方法的过度依赖,这种方法有风险——在某些情况下已经导致了——虚幻的进展和误导性的结论。我们确定了一系列限制,我们认为这些限制阻碍了可解释性研究的有意义的进展,并检查了两类流行的可解释性方法-显著性和基于单神经元的方法-作为过度依赖直觉和缺乏可证伪性如何破坏可解释性研究的案例研究。为了解决这些问题,我们提出了一种策略,以强有力的可证伪性可解释性研究框架的形式来解决这些障碍。我们鼓励研究人员利用他们的直觉作为起点,开发和测试清晰的、可证伪的假设,并希望我们的框架产生强大的、基于证据的可解释性方法,从而在我们对dnn的理解中产生有意义的进展。
Methods for understanding the decisions of and mechanisms underlying deep neural networks (DNNs) typically rely on building intuition by emphasizing sensory or semantic features of individual examples. For instance, methods aim to visualize the components of an input which are "important" to a network's decision, or to measure the semantic properties of single neurons. Here, we argue that interpretability research suffers from an over-reliance on intuition-based approaches that risk-and in some cases have caused-illusory progress and misleading conclusions. We identify a set of limitations that we argue impede meaningful progress in interpretability research, and examine two popular classes of interpretability methods-saliency and single-neuron-based approaches-that serve as case studies for how overreliance on intuition and lack of falsifiability can undermine interpretability research. To address these concerns, we propose a strategy to address these impediments in the form of a framework for strongly falsifiable interpretability research. We encourage researchers to use their intuitions as a starting point to develop and test clear, falsifiable hypotheses, and hope that our framework yields robust, evidence-based interpretability methods that generate meaningful advances in our understanding of DNNs.