Psychological models of human and optimal performance in bandit problems

Psychological models of human and optimal performance in bandit problems
复制标题

DOI:
10.1016/j.cogsys.2010.07.007
复制
发表时间:
2011-06-01
影响因子:
3.9
通讯作者:
Steyvers, Mark
Steyvers, Mark
中科院分区:
心理学3区
文献类型:
--
作者:
Lee, Michael D.;Zhang, Shunan;Steyvers, Mark

文献摘要

被引文献

相似文献

在强盗问题中,决策者必须在一组备选方案中进行选择,每个备选方案都有固定但未知的奖励率,以在一系列试验中最大化其奖励总数。要在这些问题上表现良好,需要平衡寻找高回报替代方案的需要与利用那些已知相当好的替代方案的需要。与这一动机相一致,我们开发了一种新的心理模型,该模型依赖于潜在探索和利用状态之间的切换。我们针对人类和最优决策数据,在一系列二选一强盗问题上测试该模型,并将其与强化学习文献中的基准模型进行比较。通过从最佳决策行为中推断潜在状态,我们描述了人们应该如何在探索和利用之间切换。通过从人类数据中进行推断,我们开始描述人们实际上是如何进行转换的。我们讨论了这些发现对于理解和衡量顺序决策中勘探和开发的竞争需求的影响。 (C) 2010 Elsevier B.V. 保留所有权利。
In bandit problems, a decision-maker must choose between a set of alternatives, each of which has a fixed but unknown rate of reward, to maximize their total number of rewards over a sequence of trials. Performing well in these problems requires balancing the need to search for highly-rewarding alternatives, with the need to capitalize on those alternatives already known to be reasonably good. Consistent with this motivation, we develop a new psychological model that relies on switching between latent exploration and exploitation states. We test the model over a range of two-alternative bandit problems, against both human and optimal decision-making data, comparing it to benchmark models from the reinforcement learning literature. By making inferences about the latent states from optimal decision-making behavior, we characterize how people should switch between exploration and exploitation. By making inferences from human data, we begin to characterize how people actually do switch. We discuss the implications of these findings for understanding and measuring the competing demands of exploration and exploitation in sequential decision-making. (C) 2010 Elsevier B.V. All rights reserved.