Human-AI Learning Performance in Multi-Armed Bandits

Human-AI Learning Performance in Multi-Armed Bandits
复制标题

多臂强盗中的人类人工智能学习表现

DOI:
10.1145/3306618.3314245
复制
发表时间:
2019
期刊:
Ethics and Society (AIES
影响因子:
--
通讯作者:
Dragan, Anca D.
Dragan, Anca D.
中科院分区:
--
文献类型:
--
作者:
Pandya, Ravi;Huang, Sandy H.;Hadfield-Menell, Dylan;Dragan, Anca D.

文献摘要

参考文献

被引文献

相似文献

人们经常面临具有挑战性的决策问题,其中结果不确定或未知。人工智能(AI)算法的存在可以在学习这些任务方面胜过人类。因此,AI代理有机会帮助人们更有效地学习这些任务。在这项工作中,我们使用一个多臂强盗作为一个控制设置,在其中探索这个方向。我们将人类与选定的代理配对,并观察每个人类代理团队的表现。我们发现,团队的表现可以击败人类和代理人的表现孤立。有趣的是,我们还发现,一个代理人的表现在隔离并不一定与人类代理团队的表现相关。座席性能的下降可能会导致团队性能的不成比例的大幅下降,或者在某些设置中甚至可以提高团队性能。将一个人与一个表现稍好的代理配对可以使他们表现得更好,而将他们与一个表现相同的代理配对可以使他们表现得更差。此外,我们的研究结果表明,人们有不同的探索策略,并可能表现得更好的代理,以配合他们的策略。总的来说,优化人类代理团队绩效需要超越优化代理绩效,了解代理的建议将如何影响人类决策。
People frequently face challenging decision-making problems in which outcomes are uncertain or unknown. Artificial intelligence (AI) algorithms exist that can outperform humans at learning such tasks. Thus, there is an opportunity for AI agents to assist people in learning these tasks more effectively. In this work, we use a multi-armed bandit as a controlled setting in which to explore this direction. We pair humans with a selection of agents and observe how well each human-agent team performs. We find that team performance can beat both human and agent performance in isolation. Interestingly, we also find that an agent's performance in isolation does not necessarily correlate with the human-agent team's performance. A drop in agent performance can lead to a disproportionately large drop in team performance, or in some settings can even improve team performance. Pairing a human with an agent that performs slightly better than them can make them perform much better, while pairing them with an agent that performs the same can make them them perform much worse. Further, our results suggest that people have different exploration strategies and might perform better with agents that match their strategy. Overall, optimizing human-agent team performance requires going beyond optimizing agent performance, to understanding how the agent's suggestions will influence human decision-making.
受人类启发的搜索算法 人机多臂老虎机问题的框架
DOI: --
发表时间: 2014
期刊:
影响因子: --
作者:
Paul B. Reverdy
通讯作者: Paul B. Reverdy
DOI: 10.1109/hri.2013.6483499
发表时间: 2013-03
期刊: 2013 8th ACM/IEEE International Conference on Human-Robot Interaction (HRI)
影响因子: --
作者:
S. Nikolaidis;J. Shah
通讯作者: S. Nikolaidis;J. Shah
DOI: 10.1007/s001990050146
发表时间: 1997-06-01
期刊: ECONOMIC THEORY
影响因子: 1.3
作者:
Banks, J;Olson, M;Porter, D
通讯作者: Porter, D
DOI: 10.1287/mnsc.41.5.817
发表时间: 1995-05-01
期刊: MANAGEMENT SCIENCE
影响因子: 5.4
作者:
MEYER, RJ;SHI, Y
通讯作者: SHI, Y
多臂老虎机问题策略的行为模型
DOI: 10.7907/94b0-zb90
发表时间: 2001
影响因子: 5
作者:
Christopher M. Anderson
通讯作者: Christopher M. Anderson