Principles and Practice of Explainable Machine Learning.

Principles and Practice of Explainable Machine Learning.
复制标题

DOI:
10.3389/fdata.2021.688969
复制
发表时间:
2021
影响因子:
3.1
通讯作者:
Papantonis I
Papantonis I
中科院分区:
其他
文献类型:
--
作者:
Belle V;Papantonis I

文献摘要

参考文献

被引文献

相似文献

人工智能(AI)为改善私人和公共生活提供了许多机会。自动发现海量数据中的模式和结构是数据科学的核心组成部分,目前推动着计算生物学、法律和金融等不同领域的应用。然而,这种高度积极的影响伴随着一个重大挑战:我们如何理解这些系统所建议的决策,以便我们能够信任它们?在本报告中,我们特别关注数据驱动的方法-特别是机器学习(ML)和模式识别模型-以便从文献中调查和提取结果和观察结果。本报告的目的尤其值得注意的是,ML模型正越来越多地部署在广泛的业务中。然而,随着方法的日益流行和复杂性,业务利益相关者至少对模型的缺陷、特定于数据的偏差等有越来越多的担忧。类似地,数据科学从业者通常不知道学术文献中出现的方法,或者可能难以理解不同方法之间的差异,因此最终使用Shap等行业标准。在这里,我们进行了一项调查,以帮助行业从业者(以及更广泛的数据科学家)更好地理解可解释的机器学习领域,并应用正确的工具。我们后面的部分围绕一位假定的数据科学家展开叙事,并讨论她可能如何通过提出正确的问题来解释她的模型。从组织的角度来看,在广泛地激励了该领域之后,我们讨论了主要的发展,包括允许我们研究透明模型与不透明模型的原则,以及特定于模型或与模型无关的后自组织可解释性方法。我们还对深度学习模型进行了简要的反思,并对未来的研究方向进行了讨论。
Artificial intelligence (AI) provides many opportunities to improve private and public life. Discovering patterns and structures in large troves of data in an automated manner is a core component of data science, and currently drives applications in diverse areas such as computational biology, law and finance. However, such a highly positive impact is coupled with a significant challenge: how do we understand the decisions suggested by these systems in order that we can trust them? In this report, we focus specifically on data-driven methods—machine learning (ML) and pattern recognition models in particular—so as to survey and distill the results and observations from the literature. The purpose of this report can be especially appreciated by noting that ML models are increasingly deployed in a wide range of businesses. However, with the increasing prevalence and complexity of methods, business stakeholders in the very least have a growing number of concerns about the drawbacks of models, data-specific biases, and so on. Analogously, data science practitioners are often not aware about approaches emerging from the academic literature or may struggle to appreciate the differences between different methods, so end up using industry standards such as SHAP. Here, we have undertaken a survey to help industry practitioners (but also data scientists more broadly) understand the field of explainable machine learning better and apply the right tools. Our latter sections build a narrative around a putative data scientist, and discuss how she might go about explaining her models by asking the right questions. From an organization viewpoint, after motivating the area broadly, we discuss the main developments, including the principles that allow us to study transparent models vs. opaque models, as well as model-specific or model-agnostic post-hoc explainability approaches. We also briefly reflect on deep learning models, and conclude with a discussion about future research directions.
DOI: 10.1007/s11063-011-9207-8
发表时间: 2012-04-01
影响因子: 3.1
作者:
Augasta, M. Gethsiyal;Kathirvalavakumar, T.
通讯作者: Kathirvalavakumar, T.
DOI: 10.1016/j.mineng.2012.05.008
发表时间: 2012-08-01
影响因子: 4.8
作者:
Auret, Lidia;Aldrich, Chris
通讯作者: Aldrich, Chris
DOI: 10.1007/bf00994018
发表时间: 1995-09-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
CORTES, C;VAPNIK, V
通讯作者: VAPNIK, V
DOI: 10.1162/15324430260185565
发表时间: 2002-03-01
影响因子: 6
作者:
Ben-Hur, A;Horn, D;Vapnik, V
通讯作者: Vapnik, V
DOI: 10.2307/1268249
发表时间: 1977-01-01
期刊: TECHNOMETRICS
影响因子: 2.5
作者:
COOK, RD
通讯作者: COOK, RD