Sibyl: Understanding and Addressing the Usability Challenges of Machine Learning In High-Stakes Decision Making

Sibyl: Understanding and Addressing the Usability Challenges of Machine Learning In High-Stakes Decision Making
复制标题

Sibyl:理解和解决机器学习在高风险决策中的可用性挑战

DOI:
--
复制
发表时间:
2021
影响因子:
5.2
通讯作者:
K. Veeramachaneni
K. Veeramachaneni
中科院分区:
计算机科学1区
文献类型:
--
作者:
Alexandra Zytek;Dongyu Liu;R. Vaithianathan;K. Veeramachaneni

文献摘要

参考文献

被引文献

相似文献

机器学习 (ML) 正在应用于各种且不断增长的领域。在许多情况下,领域专家(通常没有机器学习或数据科学方面的专业知识)被要求使用机器学习预测来做出高风险决策。结果可能会出现多种 ML 可用性挑战,例如用户对模型缺乏信任、无法调和人类与 ML 的分歧,以及对将复杂问题过度简化为单一算法输出的道德担忧。在本文中,我们通过与儿童福利筛查人员的一系列合作,调查了儿童福利筛查领域中存在的机器学习可用性挑战。在机器学习科学家、可视化研究人员和领域专家(儿童筛查人员)之间的迭代设计过程之后,我们首先确定了四个关键的机器学习挑战,并磨练了一种有前景的可解释机器学习技术来解决这些问题(局部因素贡献)。然后,我们实施并评估了我们的可视化分析工具 SIBYL,以提高本地因素贡献的可解释性和交互性。两项正式用户研究分别由 12 名非专家参与者和 13 名专家参与者证明了我们工具的有效性。我们收集了宝贵的反馈,从中列出了一系列设计含义,作为研究人员的有用指南,这些研究人员旨在为儿童福利筛查人员和其他类似领域专家部署的 ML 预测模型开发可解释的交互式可视化工具。
Machine learning (ML) is being applied to a diverse and ever-growing set of domains. In many cases, domain experts – who often have no expertise in ML or data science – are asked to use ML predictions to make high-stakes decisions. Multiple ML usability challenges can appear as result, such as lack of user trust in the model, inability to reconcile human-ML disagreement, and ethical concerns about oversimplification of complex problems to a single algorithm output. In this paper, we investigate the ML usability challenges that present in the domain of child welfare screening through a series of collaborations with child welfare screeners. Following the iterative design process between the ML scientists, visualization researchers, and domain experts (child screeners), we first identified four key ML challenges and honed in on one promising explainable ML technique to address them (local factor contributions). Then we implemented and evaluated our visual analytics tool, SIBYL, to increase the interpretability and interactivity of local factor contributions. The effectiveness of our tool is demonstrated by two formal user studies with 12 non-expert participants and 13 expert participants respectively. Valuable feedback was collected, from which we composed a list of design implications as a useful guideline for researchers who aim to develop an interpretable and interactive visualization tool for ML prediction models deployed for child welfare screeners and other similar domain experts.
DOI: 10.1016/j.jmrt.2023.02.141
发表时间: 2023-03-04
影响因子: 6.4
作者:
Wang, Qingjuan;He, Zeen;Yang, Congcong
通讯作者: Yang, Congcong