Finding AI’s Faults with AAR/AI: An Empirical Study

Finding AI’s Faults with AAR/AI: An Empirical Study
复制标题

DOI:
10.1145/3487065
复制
发表时间:
2022-03
期刊:
ACM Transactions on Interactive Intelligent Systems (TiiS)
影响因子:
--
通讯作者:
Roli Khanna;Jonathan Dodge;Andrew Anderson;Rupika Dikkala;Jed Irvine;Zeyad Shureih;Kin-Ho Lam;Caleb R. Matthews;Zhengxian Lin;Minsuk Kahng;Alan Fern;M. Burnett
Roli Khanna;Jonathan Dodge;Andrew Anderson;Rupika Dikkala;Jed Irvine;Zeyad Shureih;Kin-Ho Lam;Caleb R. Matthews;Zhengxian Lin;Minsuk Kahng;Alan Fern;M. Burnett
中科院分区:
其他
文献类型:
--
作者:
Roli Khanna;Jonathan Dodge;Andrew Anderson;Rupika Dikkala;Jed Irvine;Zeyad Shureih;Kin-Ho Lam;Caleb R. Matthews;Zhengxian Lin;Minsuk Kahng;Alan Fern;M. Burnett

文献摘要

被引文献

相似文献

你会允许AI代理代表你做决定吗?如果答案是“不总是”,那么下一个问题就变成了“在什么情况下”?回答这个问题需要人类用户能够评估AI代理-而不仅仅是整体通过/失败评估或统计数据。在这里,用户需要能够定位代理的错误,以便他们可以确定何时愿意依赖代理,何时不愿意。人工智能事后评估(AAR/AI)是一种新的人工智能评估过程,用于与可解释的人工智能系统集成,旨在支持人类用户在这奋进的努力,在本文中,我们实证研究了AAR/AI对领域知识丰富的用户的有效性。我们的研究结果表明,AAR/AI参与者不仅比非AAR/AI参与者发现了更多的错误(即,显示出更大的回忆)而且更精确地定位它们(即,更精确)。事实上,AAR/AI参与者在每个错误上的表现都优于非AAR/AI参与者,平均而言,发现任何特定错误的可能性几乎是非AAR/AI参与者的六倍。最后,有证据表明,将标签纳入AAR/AI过程可能会鼓励领域知识丰富的用户抽象上述个别错误实例;我们假设这样做可能会进一步提高AAR/AI参与者的有效性。
Would you allow an AI agent to make decisions on your behalf? If the answer is “not always,” the next question becomes “in what circumstances”? Answering this question requires human users to be able to assess an AI agent—and not just with overall pass/fail assessments or statistics. Here users need to be able to localize an agent’s bugs so that they can determine when they are willing to rely on the agent and when they are not. After-Action Review for AI (AAR/AI), a new AI assessment process for integration with Explainable AI systems, aims to support human users in this endeavor, and in this article we empirically investigate AAR/AI’s effectiveness with domain-knowledgeable users. Our results show that AAR/AI participants not only located significantly more bugs than non-AAR/AI participants did (i.e., showed greater recall) but also located them more precisely (i.e., with greater precision). In fact, AAR/AI participants outperformed non-AAR/AI participants on every bug and were, on average, almost six times as likely as non-AAR/AI participants to find any particular bug. Finally, evidence suggests that incorporating labeling into the AAR/AI process may encourage domain-knowledgeable users to abstract above individual instances of bugs; we hypothesize that doing so may have contributed further to AAR/AI participants’ effectiveness.