Evaluating effects of automation reliability and reliability information on trust, dependence and dual-task performance

Evaluating effects of automation reliability and reliability information on trust, dependence and dual-task performance
复制标题

评估自动化可靠性和可靠性信息对信任、依赖和双任务性能的影响

DOI:
--
复制
发表时间:
2018
期刊:
Proceedings of the Human Factors and Ergonomics Society Annual Meeting
影响因子:
--
通讯作者:
Na Du
Na Du
中科院分区:
--
文献类型:
--
作者:
Na Du

文献摘要

被引文献

相似文献

使用自动决策辅助工具可以减少人类面临的危险,并使人类工作人员能够执行更具挑战性的任务。然而,当人们不能适当地信任和依赖自动化时,自动化就会出现问题。现有的研究已经表明,为用户提供包括自动化确定性、可靠性和置信度的可能性信息的系统设计可以促进信任-可靠性校准、人对自动化的信任与自动化的能力之间的对应(Lee & Moray,1994),并提高人类自动化任务性能(Beller等人,2013; Wang,Jamieson,& Hollands,2009; McGuirl & Sarter,2006)。虽然已经提出了披露可靠性信息作为设计解决方案,但这种信息披露的具体效果仍然各不相同(Wang等人,2009;弗莱彻等人,2017; Walliser等人,2016年)。明确的指导方针,将允许显示设计师选择最有效的可靠性信息,以促进人类的决策性能和信任校准似乎不存在。因此,本研究的目的是调和现有的文献调查,如果以及如何不同的方法计算可靠性信息影响其有效性在不同的自动化可靠性。一项人类受试者实验由60名参与者进行。每个参与者进行了补偿跟踪任务和威胁检测任务,同时与一个不完美的自动威胁检测器的帮助下。实验采用2×4混合设计,两个自变量为:自动化可靠性(68%vs.90%)作为被试内因素,可靠性信息作为被试间因素。基于信号检测理论和贝叶斯定理的条件概率公式(H:命中; CR:正确拒绝; FA:虚警; M:未命中),使用不同方法计算自动威胁检测器的可靠性信息:总体可靠性= P(H + CR| H + FA + M + CR)。阳性预测值= P(H|阴性预测值= P(CR| CR + M)。命中率= P(H| H + M),正确拒绝率= P(CR| CR + FA)。还有一个控制条件,即参与者没有被告知任何可靠性信息,而只是被告知自动威胁检测器发出的警报可能正确,也可能不正确。感兴趣的因变量是参与者对自动化的主观信任和对他们的显示切换行为的客观测量。这项研究的结果表明,随着自动威胁检测器变得更加可靠,参与者对威胁检测器的信任和依赖程度显着增加,他们的检测性能也有所提高。更重要的是,当信度信息采用不同的计算方法时,被试的信任、依赖和双任务绩效存在显著差异。具体来说,当自动威胁检测器的整体可靠性为90%时,揭示自动化的积极和消极预测值显着帮助参与者校准他们对检测器的信任和依赖,并导致检测任务的最短反应时间。然而,当自动威胁检测器的总体可靠性为68%时,阳性和阴性预测值并没有导致参与者对检测器的依从性有显著差异。此外,我们的研究结果表明,命中率和正确拒绝率或整体可靠性的披露似乎并没有帮助人类自动化团队的性能和信任可靠性校准。这项研究的一个意义是,用户应该意识到系统的可靠性,特别是积极/消极的预测值,产生适当的信任和依赖的自动化。这可以应用于自动决策辅助的界面设计。未来的研究应该研究当自动威胁检测器的标准变得自由时,阳性和阴性预测值是否仍然是用于信任校准的最有效的信息。
The use of automated decision aids could reduce human exposure to dangers and enable human workers to perform more challenging tasks. However, automation is problematic when people fail to trust and depend on it appropriately. Existing studies have shown that system design that provides users with likelihood information including automation certainty, reliability, and confidence could facilitate trust- reliability calibration, the correspondence between a person’s trust in the automation and the automation’s capabilities (Lee & Moray, 1994), and improve human–automation task performance (Beller et al., 2013; Wang, Jamieson, & Hollands, 2009; McGuirl & Sarter, 2006). While revealing reliability information has been proposed as a design solution, the concrete effects of such information disclosure still vary (Wang et al., 2009; Fletcher et al., 2017; Walliser et al., 2016). Clear guidelines that would allow display designers to choose the most effective reliability information to facilitate human decision performance and trust calibration do not appear to exist. The present study, therefore, aimed to reconcile existing literature by investigating if and how different methods of calculating reliability information affect their effectiveness at different automation reliability. A human subject experiment was conducted with 60 participants. Each participant performed a compensatory tracking task and a threat detection task simultaneously with the help of an imperfect automated threat detector. The experiment adopted a 2×4 mixed design with two independent variables: automation reliability (68% vs. 90%) as a within- subject factor and reliability information as a between-subjects factor. Reliability information of the automated threat detector was calculated using different methods based on the signal detection theory and conditional probability formula of Bayes’ Theorem (H: hits; CR: correct rejections, FA: false alarms; M: misses): Overall reliability = P (H + CR | H + FA + M + CR). Positive predictive value = P (H | H + FA); negative predictive value = P (CR | CR + M). Hit rate = P (H | H + M), correct rejection rate = P (CR | CR + FA). There was also a control condition where participants were not informed of any reliability information but only told the alerts from the automated threat detector may or may not be correct. The dependent variables of interest were participants’ subjective trust in automation and objective measures of their display-switching behaviors. The results of this study showed that as the automated threat detector became more reliable, participants’ trust in and dependence on the threat detector increased significantly, and their detection performance improved. More importantly, there were significant differences in participants’ trust, dependence and dual-task performance when reliability information was calculated by different methods. Specifically, when overall reliability of the automated threat detector was 90%, revealing positive and negative predictive values of the automation significantly helped participants to calibrate their trust in and dependence on the detector, and led to the shortest reaction time for detection task. However, when overall reliability of the automated threat detector was 68%, positive and negative predictive values didn’t lead to significant difference in participants’ compliance on the detector. In addition, our result demonstrated that the disclosure of hit rate and correct rejection rate or overall reliability didn’t seem to aid human-automation team performance and trust-reliability calibration. An implication of the study is that users should be made aware of system reliability, especially of positive/negative predictive values, to engender appropriate trust in and dependence on the automation. This can be applied to the interface design of automated decision aids. Future studies should examine whether the positive and negative predictive values are still the most effective pieces of information for trust calibration when the criterion of the automated threat detector becomes liberal.