From local explanations to global understanding with explainable AI for trees

From local explanations to global understanding with explainable AI for trees
复制标题

DOI:
10.1038/s42256-019-0138-9
复制
发表时间:
2020-01-01
影响因子:
23.8
通讯作者:
Lee, Su-In
Lee, Su-In
中科院分区:
计算机科学1区
文献类型:
--
作者:
Lundberg, Scott M.;Erion, Gabriel;Lee, Su-In

文献摘要

被引文献

相似文献

基于树的机器学习模型,如随机森林、决策树和梯度提升树是流行的非线性预测模型,但相对较少关注解释它们的预测。在这里,我们通过三个主要贡献来提高基于树的模型的可解释性。(1)一个基于博弈论的计算最优解释的多项式时间算法。(2)一种直接测量局部特征相互作用效应的新型解释。(3)一套新的工具,用于理解基于结合每个预测的许多局部解释的全局模型结构。我们将这些工具应用于三个医疗机器学习问题,并展示了如何结合许多高质量的局部解释,使我们能够表示全局结构,同时保持对原始模型的局部忠实性。这些工具使我们能够(1)识别美国人群中高幅度但低频率的非线性死亡风险因素,(2)突出具有共同风险特征的不同人群亚组,(3)识别慢性肾脏疾病风险因素之间的非线性相互作用效应,以及(4)通过识别哪些特征会随着时间的推移降低模型的性能来监控医院中部署的机器学习模型。鉴于基于树的机器学习模型的流行,这些对其可解释性的改进在广泛的领域中具有影响。基于树的机器学习模型广泛应用于医疗保健、金融和公共服务等领域。作者提出了一种树的解释方法,该方法可以计算个体预测的最佳局部解释,并在三个医学数据集上演示了他们的方法。
Tree-based machine learning models such as random forests, decision trees and gradient boosted trees are popular nonlinear predictive models, yet comparatively little attention has been paid to explaining their predictions. Here we improve the interpretability of tree-based models through three main contributions. (1) A polynomial time algorithm to compute optimal explanations based on game theory. (2) A new type of explanation that directly measures local feature interaction effects. (3) A new set of tools for understanding global model structure based on combining many local explanations of each prediction. We apply these tools to three medical machine learning problems and show how combining many high-quality local explanations allows us to represent global structure while retaining local faithfulness to the original model. These tools enable us to (1) identify high-magnitude but low-frequency nonlinear mortality risk factors in the US population, (2) highlight distinct population subgroups with shared risk characteristics, (3) identify nonlinear interaction effects among risk factors for chronic kidney disease and (4) monitor a machine learning model deployed in a hospital by identifying which features are degrading the model's performance over time. Given the popularity of tree-based machine learning models, these improvements to their interpretability have implications across a broad set of domains. Tree-based machine learning models are widely used in domains such as healthcare, finance and public services. The authors present an explanation method for trees that enables the computation of optimal local explanations for individual predictions, and demonstrate their method on three medical datasets.