Humans use directed and random exploration to solve the explore-exploit dilemma.

Humans use directed and random exploration to solve the explore-exploit dilemma.
复制标题

人类使用定向和随机探索来解决探索-利用困境。

DOI:
10.1037/a0038199
复制
发表时间:
2014
期刊:
Journal of experimental psychology. General
影响因子:
--
通讯作者:
Cohen,JonathanD
Cohen,JonathanD
中科院分区:
--
文献类型:
--
作者:
Wilson,RobertC;Geana,Andra;White,JohnM;Ludvig,ElliotA;Cohen,JonathanD

文献摘要

参考文献

被引文献

相似文献

所有具有适应性的生物都面临着一个基本的权衡:是追求已知的奖励(开发),还是为了寻找更好的东西(探索)而取样不太为人所知的选择。理论建议至少有两种策略来解决这一困境:一种是定向策略,其中选择明显偏向于信息寻求,另一种是随机策略,其中决策噪音导致偶然的探索。在这项工作中,我们调查了人类使用这两种策略的程度。在我们的“视界任务”中,参与者在两种不同的情况下做出探索-利用决策,这两种情况下他们在未来(时间范围)将做出的选择数量不同。参与者可以在每个游戏中做出一个选择(视界1),或者连续做出6个选择(视界6),这给了他们更多的探索机会。通过对这两种情况下的行为进行建模,我们能够测量与探索相关的决策变化,并量化这两种策略对行为的贡献。研究发现,在较长的视界范围内,参与者更倾向于寻求信息,并且有更高的决策噪声,这表明人类使用这两种策略来解决探索-开发困境。因此,我们得出的结论是,信息寻求和选择可变性都可以被控制,并用于勘探服务。(PsycINFO数据库记录(c) 2016 APA,版权所有)
All adaptive organisms face the fundamental tradeoff between pursuing a known reward (exploitation) and sampling lesser-known options in search of something better (exploration). Theory suggests at least two strategies for solving this dilemma: a directed strategy in which choices are explicitly biased toward information seeking, and a random strategy in which decision noise leads to exploration by chance. In this work we investigated the extent to which humans use these two strategies. In our “Horizon task,” participants made explore–exploit decisions in two contexts that differed in the number of choices that they would make in the future (the time horizon). Participants were allowed to make either a single choice in each game (horizon 1), or 6 sequential choices (horizon 6), giving them more opportunity to explore. By modeling the behavior in these two conditions, we were able to measure exploration-related changes in decision making and quantify the contributions of the two strategies to behavior. We found that participants were more information seeking and had higher decision noise with the longer horizon, suggesting that humans use both strategies to solve the exploration–exploitation dilemma. We thus conclude that both information seeking and choice variability can be controlled and put to use in the service of exploration.(PsycINFO Database Record (c) 2016 APA, all rights reserved)
DOI: 10.2307/1268092
发表时间: 1975
期刊: Technometrics
影响因子: 2.5
作者:
J. Gani;K. Sarkadi;I. Vincze
通讯作者: I. Vincze
一种新颖的鸟鸣发声学习强化模型
DOI: --
发表时间: 1994
期刊: Neural Information Processing Systems
影响因子: --
作者:
K. Doya;T. Sejnowski
通讯作者: T. Sejnowski
DOI: 10.1016/j.cogsys.2010.07.007
发表时间: 2011-06-01
影响因子: 3.9
作者:
Lee, Michael D.;Zhang, Shunan;Steyvers, Mark
通讯作者: Steyvers, Mark
DOI: 10.3389/fnins.2012.00150
发表时间: 2012
影响因子: 4.3
作者:
Payzan-Lenestour E;Bossaerts P
通讯作者: Bossaerts P
DOI: 10.1002/bdm.3960070303
发表时间: 1994-09-01
影响因子: 2
作者:
BIER, VM;CONNELL, BL
通讯作者: CONNELL, BL