Opening the black box of artificial intelligence for clinical decision support: A study predicting stroke outcome

Opening the black box of artificial intelligence for clinical decision support: A study predicting stroke outcome
复制标题

DOI:
10.1371/journal.pone.0231166
复制
发表时间:
2020-04-06
期刊:
影响因子:
3.7
通讯作者:
Frey, Dietmar
Frey, Dietmar
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Zihni, Esra;Madai, Vince Istvan;Frey, Dietmar

文献摘要

被引文献

相似文献

最先进的机器学习 (ML) 人工智能方法越来越多地应用于临床预测模型,为医生提供临床决策支持系统。人工神经网络 (ANN) 和树提升等现代机器学习方法通​​常比逻辑回归等传统方法表现更好。另一方面,这些现代方法对最终预测的理解有限。然而,在医学领域,对应用模型的理解至关重要,特别是在为临床决策支持提供信息时。因此,近年来,现代机器学习方法的可解释性方法已经出现,有可能实现可解释的预测与高性能的结合。据我们所知,我们在这项工作中首次展示了两种现代机器学习方法(树增强和多层感知器(MLP))与使用中风结果预测范式的传统逻辑回归方法的可解释性比较。在这里,我们使用临床特征来预测二分法的中风后 90 天改良 Rankin 量表 (mRS) 评分。为了可解释性,我们使用 MLP 的深度泰勒分解、树增强的 Shapley 值和逻辑回归的模型系数来评估临床特征的重要性。就测试数据集上的曲线下面积 (AUC) 值衡量的性能而言,所有模型的表现相当:三种不同正则化方案的逻辑回归 AUC 分别为 0.83、0.83、0.81;树增强 AUC 为 0.81; MLP AUC 为 0.83。重要的是,可解释性分析通过将年龄和中风严重程度连续评为最重要的预测特征,证明了跨模型的结果一致。对于不太重要的特征,在方法之间观察到了一些差异。我们的分析表明,现代机器学习方法可以提供与领域知识解释和传统方法排名兼容的可解释性。未来的工作应侧重于在其他数据集中复制这些发现,并进一步测试不同的可解释性方法。
State-of-the-art machine learning (ML) artificial intelligence methods are increasingly lever-aged in clinical predictive modeling to provide clinical decision support systems to physicians. Modern ML approaches such as artificial neural networks (ANNs) and tree boosting often perform better than more traditional methods like logistic regression. On the other hand, these modern methods yield a limited understanding of the resulting predictions. However, in the medical domain, understanding of applied models is essential, in particular, when informing clinical decision support. Thus, in recent years, interpretability methods for modern ML methods have emerged to potentially allow explainable predictions paired with high performance. To our knowledge, we present in this work the first explainability comparison of two modern ML methods, tree boosting and multilayer perceptrons (MLPs), to traditional logistic regression methods using a stroke outcome prediction paradigm. Here, we used clinical features to predict a dichotomized 90 days post-stroke modified Rankin Scale (mRS) score. For interpretability, we evaluated clinical features' importance with regard to predictions using deep Taylor decomposition for MLP, Shapley values for tree boosting and model coefficients for logistic regression. With regard to performance as measured by Area under the Curve (AUC) values on the test dataset, all models performed comparably: Logistic regression AUCs were 0.83, 0.83, 0.81 for three different regularization schemes; tree boosting AUC was 0.81; MLP AUC was 0.83. Importantly, the interpretability analysis demonstrated consistent results across models by rating age and stroke severity consecutively amongst the most important predictive features. For less important features, some differences were observed between the methods. Our analysis suggests that modern machine learning methods can provide explainability which is compatible with domain knowledge interpretation and traditional method rankings. Future work should focus on replication of these findings in other datasets and further testing of different explainability methods.