Exploration and Exploitation During Sequential Search

Exploration and Exploitation During Sequential Search
复制标题

DOI:
10.1111/j.1551-6709.2009.01021.x
复制
发表时间:
2009-05-01
期刊:
影响因子:
2.5
通讯作者:
Koerding, Konrad
Koerding, Konrad
中科院分区:
心理学3区
文献类型:
--
作者:
Dam, Gregory;Koerding, Konrad

文献摘要

被引文献

相似文献

当我们学习如何投掷飞镖时,我们调整了如何将基础油投掷到飞镖粘附的地方。许多技能学习在计算上是相似的,因为我们是通过完成单个动作后获得的反馈来学习的。我们可以将这类任务形式化为在所有可能的行动中寻找能带来最高回报的行动。在这种情况下,我们的行动有两个目标:我们想要最好地利用我们已经知道的(开发),但我们也想要学习在未来更成功(探索)。在这里,我们测试了参与者是如何学习运动轨迹的,反馈是根据所选择的轨迹提供的金钱奖励。我们用决策理论从数学上推导出实验的最优搜索策略。一个理想的搜索者模型可以很好地预测参与者的搜索行为,该模型将探索和利用相结合。
When we learn how to throw darts we adjust how we throw based oil where the darts stick. Much of skill learning is computationally similar in that we learn using feedback obtained after the completion of individual actions. We can formalize such tasks as a search problem among the set of all possible actions, find the action that leads to the highest reward. In such cases our actions have two objectives: we want to best utilize what we already know (exploitation), but we also want to learn to be more successful in the future (exploration). Here we tested how participants learn movement trajectories where feedback is provided as a monetary reward that depends on the chosen trajectory. We mathematically derived the optimal search policy for our experiment using decision theory. The search behavior of participants is well predicted by an ideal searcher model that optimally combines exploration and exploitation.