Learning the value of information and reward over time when solving exploration-exploitation problems

Learning the value of information and reward over time when solving exploration-exploitation problems
复制标题

DOI:
10.1038/s41598-017-17237-w
复制
发表时间:
2017-12-05
期刊:
影响因子:
4.6
通讯作者:
Alexander, William
Alexander, William
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Dezza, Irene Cogliati;Yu, Angela J.;Alexander, William

文献摘要

被引文献

相似文献

为了灵活地适应环境的要求,动物们不断地暴露在冲突中,这是因为它们不得不在可预测的奖励熟悉的选择(剥削)和有风险的新选择之间做出选择,其价值基本上包括获得关于可能奖励空间的新信息(探索)。尽管进行了广泛的研究,但动物解决这种剥削-探索困境的机制仍然知之甚少。在这里,我们调查人类决策的赌博任务中,每个审判的信息价值和奖励潜力分别操纵。为了更好地描述强调观察到的行为选择的机制,我们引入了一个计算模型,通过将值与信息相关联来增强标准的基于奖励的强化学习公式。我们发现,在学习过程中获得的奖励和信息的影响之间的平衡利用和探索,这种影响是依赖于奖励的背景。我们的研究结果揭示了在不确定性下决策的机制,并提出了新的方法来调查整个动物王国的探索-开发困境。
To flexibly adapt to the demands of their environment, animals are constantly exposed to the conflict resulting from having to choose between predictably rewarding familiar options (exploitation) and risky novel options, the value of which essentially consists of obtaining new information about the space of possible rewards (exploration). Despite extensive research, the mechanisms that subtend the manner in which animals solve this exploitation-exploration dilemma are still poorly understood. Here, we investigate human decision-making in a gambling task in which the informational value of each trial and the reward potential were separately manipulated. To better characterize the mechanisms that underlined the observed behavioural choices, we introduce a computational model that augments the standard reward-based reinforcement learning formulation by associating a value to information. We find that both reward and information gained during learning influence the balance between exploitation and exploration, and that this influence was dependent on the reward context. Our results shed light on the mechanisms that underpin decision-making under uncertainty, and suggest new approaches for investigating the exploration-exploitation dilemma throughout the animal kingdom.