Designing Alternative Representations of Confusion Matrices to Support Non-Expert Public Understanding of Algorithm Performance

Designing Alternative Representations of Confusion Matrices to Support Non-Expert Public Understanding of Algorithm Performance
复制标题

DOI:
10.1145/3415224
复制
发表时间:
2020-10
影响因子:
--
通讯作者:
Hong Shen;Ángel Alexander Cabrera
Hong Shen;Ángel Alexander Cabrera
中科院分区:
--
文献类型:
--
作者:
Hong Shen;Ángel Alexander Cabrera

文献摘要

被引文献

相似文献

随着人工智能系统越来越多地部署到我们的社会中,确保公众有效理解由机器学习技术驱动的算法决策已经成为一项紧迫的任务。在这项工作中,我们通过重新设计二进制分类的混淆矩阵来支持非专家理解机器学习模型的性能,从而为实现这一目标迈出了具体的一步。通过访谈(n=7)和调查(n=102),我们绘制了两组主要的挑战,奠定了人们在理解标准混淆矩阵:一般术语和矩阵设计。我们进一步确定了关于矩阵设计的三个子挑战,即对阅读数据的方向、分层关系和所涉及的数量的混淆。然后,我们与483名参与者进行了一项在线实验,以评估一系列替代表示在预测累犯的算法背景下针对这些挑战的有效性。我们开发了三个层次的问题来评估用户的客观理解。我们评估了我们的替代方案在回答这些问题的准确性,完成时间和主观理解方面的有效性。我们的研究结果表明,(1)只有通过语境化的术语,我们可以显着提高用户的理解和(2)流程图,这有助于指出方向的阅读的数据,是最有用的,在提高客观的理解。我们的研究结果为开发更直观、更易于理解的机器学习模型性能表示奠定了基础。
Ensuring effective public understanding of algorithmic decisions that are powered by machine learning techniques has become an urgent task with the increasing deployment of AI systems into our society. In this work, we present a concrete step toward this goal by redesigning confusion matrices for binary classification to support non-experts in understanding the performance of machine learning models. Through interviews (n=7) and a survey (n=102), we mapped out two major sets of challenges lay people have in understanding standard confusion matrices: the general terminologies and the matrix design. We further identified three sub-challenges regarding the matrix design, namely, confusion about the direction of reading the data, layered relations and quantities involved. We then conducted an online experiment with 483 participants to evaluate how effective a series of alternative representations target each of those challenges in the context of an algorithm for making recidivism predictions. We developed three levels of questions to evaluate users' objective understanding. We assessed the effectiveness of our alternatives for accuracy in answering those questions, completion time, and subjective understanding. Our results suggest that (1) only by contextualizing terminologies can we significantly improve users' understanding and (2) flow charts, which help point out the direction of reading the data, were most useful in improving objective understanding. Our findings set the stage for developing more intuitive and generally understandable representations of the performance of machine learning models.