One Explanation Does Not Fit All The Promise of Interactive Explanations for Machine Learning Transparency

One Explanation Does Not Fit All The Promise of Interactive Explanations for Machine Learning Transparency
复制标题

一种解释并不能满足交互式解释对机器学习透明度的所有承诺

DOI:
10.1007/s13218-020-00637-y
复制
发表时间:
2020
期刊:
KI - Künstliche Intelligenz
影响因子:
--
通讯作者:
Sokol K
Sokol K
中科院分区:
--
文献类型:
--
作者:
Sokol K

文献摘要

相似文献

基于机器学习算法的预测系统由于其在行业中的不断扩散而产生了对透明度的需求。每当黑箱算法预测影响人类事务时,就应该仔细审查这些算法的内部工作原理,并向相关利益相关者解释它们的决定,包括系统工程师、系统操作员和正在裁决案件的个人。虽然有各种可解释性和可解释性的方法可用,但没有一种方法是灵丹妙药,可以满足有关各方可能要求的所有不同的期望和相互竞争的目标。在本文中,我们通过使用对比解释的例子来讨论交互式机器学习对于提高黑盒系统透明度的承诺来解决这一挑战-这是一种最先进的可解释机器学习方法。具体地说,我们展示了如何通过交互地调整条件语句来个性化反事实的解释,并通过询问后续的“如果呢?”来提取额外的解释。问题。我们在构建、部署和展示这类系统方面的经验使我们能够列出所需的特性以及潜在的限制,可以用来指导交互式解释程序的开发。虽然定制交互媒介,即由各种交流渠道组成的用户界面可能会给人一种个性化的印象,但我们认为,调整解释本身及其内容更重要。为此,除了明确告知被解释人其局限性和注意事项外,还必须考虑解释的广度、范围、背景、目的和目标等性质。此外,我们还讨论了映射被解释者的心理模型的挑战,这是可理解的人机交互的主要构建块。我们还考虑了允许解说者自由操纵解释从而提取关于潜在预测模型的信息的风险,这些信息可能会被恶意行为者用来窃取或玩弄模型。最后,构建端到端交互式可解释性系统是一项具有挑战性的工程任务;除非主要目标是其部署,否则我们建议将“绿野仙踪”研究作为测试和评估独立交互式可解释性算法的代理。
The need for transparency of predictive systems based on Machine Learning algorithms arises as a consequence of their ever-increasing proliferation in the industry. Whenever black-box algorithmic predictions influence human affairs, the inner workings of these algorithms should be scrutinised and their decisions explained to the relevant stakeholders, including the system engineers, the system’s operators and the individuals whose case is being decided. While a variety of interpretability and explainability methods is available, none of them is a panacea that can satisfy all diverse expectations and competing objectives that might be required by the parties involved. We address this challenge in this paper by discussing the promises ofInteractiveMachine Learning for improved transparency of black-box systems using the example of contrastive explanations—a state-of-the-art approach toInterpretableMachine Learning. Specifically, we show how to personalise counterfactual explanations by interactively adjusting their conditional statements and extract additional explanations by asking follow-up “What if?” questions. Our experience in building, deploying and presenting this type of system allowed us to list desired properties as well as potential limitations, which can be used to guide the development of interactive explainers. While customising the medium of interaction, i.e., the user interface comprising of various communication channels, may give an impression of personalisation, we argue that adjusting the explanation itself and its content is more important. To this end, properties such as breadth, scope, context, purpose and target of the explanation have to be considered, in addition to explicitly informing the explainee about its limitations and caveats. Furthermore, we discuss the challenges of mirroring the explainee’s mental model, which is the main building block of intelligible human–machine interactions. We also deliberate on the risks of allowing the explainee to freely manipulate the explanations and thereby extracting information about the underlying predictive model, which might be leveraged by malicious actors to steal or game the model. Finally, building an end-to-end interactive explainability system is a challenging engineering task; unless the main goal is its deployment, we recommend “Wizard of Oz” studies as a proxy for testing and evaluating standalone interactive explainability algorithms.