Definitions, methods, and applications in interpretable machine learning

Definitions, methods, and applications in interpretable machine learning
复制标题

DOI:
10.1073/pnas.1900654116
复制
发表时间:
2019-10-29
影响因子:
11.1
通讯作者:
Yu, Bin
Yu, Bin
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Murdoch, W. James;Singh, Chandan;Yu, Bin

文献摘要

被引文献

相似文献

机器学习模型在学习复杂模式方面取得了巨大的成功,这些模式使它们能够对未观察到的数据进行预测。除了使用模型进行预测之外,解释模型所学到的知识的能力也越来越受到关注。然而,这种日益增加的关注导致了对可解释性概念的相当大的混乱。特别是,目前尚不清楚各种拟议的解释方法之间的联系,以及可以使用哪些共同概念来评价这些方法。我们的目标是通过在机器学习的背景下定义可解释性并引入预测,描述性,相关(PDR)框架来讨论解释来解决这些问题。PDR框架为评估提供了3个首要要求:预测准确性、描述准确性和相关性,其中相关性是相对于人类受众来判断的。此外,为了帮助管理大量的解释方法,我们引入了一个分类现有的技术到基于模型的和事后类别,与子组,包括稀疏性,模块化,和模拟。为了展示从业者如何使用PDR框架来评估和理解解释,我们提供了许多真实的例子。这些例子突出了人类观众在可解释性讨论中所扮演的往往被低估的角色。最后,基于我们的框架,我们讨论了现有方法的局限性和未来的工作方向。我们希望这项工作能够提供一个通用的词汇表,使从业者和研究人员更容易讨论和选择全方位的解释方法。
Machine-learning models have demonstrated great success in learning complex patterns that enable them to make predictions about unobserved data. In addition to using models for prediction, the ability to interpret what a model has learned is receiving an increasing amount of attention. However, this increased focus has led to considerable confusion about the notion of interpretability. In particular, it is unclear how the wide array of proposed interpretation methods are related and what common concepts can be used to evaluate them. We aim to address these concerns by defining interpretability in the context of machine learning and introducing the predictive, descriptive, relevant (PDR) framework for discussing interpretations. The PDR framework provides 3 overarching desiderata for evaluation: predictive accuracy, descriptive accuracy, and relevancy, with relevancy judged relative to a human audience. Moreover, to help manage the deluge of interpretation methods, we introduce a categorization of existing techniques into model-based and post hoc categories, with subgroups including sparsity, modularity, and simulatability. To demonstrate how practitioners can use the PDR framework to evaluate and understand interpretations, we provide numerous real-world examples. These examples highlight the often underappreciated role played by human audiences in discussions of interpretability. Finally, based on our framework, we discuss limitations of existing methods and directions for future work. We hope that this work will provide a common vocabulary that will make it easier for both practitioners and researchers to discuss and choose from the full range of interpretation methods.