An Evaluation of the Human-Interpretability of Explanation

An Evaluation of the Human-Interpretability of Explanation
复制标题

DOI:
--
复制
发表时间:
2019-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Isaac Lage;Emily Chen;Jeffrey He;Menaka Narayanan;Been Kim;Sam Gershman;F. Doshi-Velez
Isaac Lage;Emily Chen;Jeffrey He;Menaka Narayanan;Been Kim;Sam Gershman;F. Doshi-Velez
中科院分区:
其他
文献类型:
--
作者:
Isaac Lage;Emily Chen;Jeffrey He;Menaka Narayanan;Been Kim;Sam Gershman;F. Doshi-Velez

文献摘要

被引文献

相似文献

近年来,人们对机器学习系统的兴趣激增,机器学习系统可以为他们的预测或决策提供人类可以理解的理由。然而,究竟什么样的解释才是真正人类可以解释的,人们仍然知之甚少。这项工作促进了我们对如何在用户可能用机器学习系统执行的三个特定任务下解释解释的理解:模拟响应,验证建议的响应,以及确定建议的响应的正确性是否在输入发生变化的情况下发生变化。通过仔细控制的人类受试者实验,我们确定了可用于优化机器学习系统的解释性的正则化规则。我们的结果表明,复杂性的类型很重要:认知块(新定义的概念)比变量重复对绩效的影响更大,而且这些趋势在任务和领域之间是一致的。这表明,解释系统可能存在一些共同的设计原则。
Recent years have seen a boom in interest in machine learning systems that can provide a human-understandable rationale for their predictions or decisions. However, exactly what kinds of explanation are truly human-interpretable remains poorly understood. This work advances our understanding of what makes explanations interpretable under three specific tasks that users may perform with machine learning systems: simulation of the response, verification of a suggested response, and determining whether the correctness of a suggested response changes under a change to the inputs. Through carefully controlled human-subject experiments, we identify regularizers that can be used to optimize for the interpretability of machine learning systems. Our results show that the type of complexity matters: cognitive chunks (newly defined concepts) affect performance more than variable repetitions, and these trends are consistent across tasks and domains. This suggests that there may exist some common design principles for explanation systems.