Neural correlates of forward planning in a spatial decision task in humans.

Neural correlates of forward planning in a spatial decision task in humans.
复制标题

DOI:
10.1523/jneurosci.4647-10.2011
复制
发表时间:
2011-04-06
期刊:
The Journal of neuroscience : the official journal of the Society for Neuroscience
影响因子:
--
通讯作者:
Daw ND
Daw ND
中科院分区:
其他
文献类型:
--
作者:
Simon DA;Daw ND

文献摘要

相似文献

虽然强化学习(RL)理论在表征大脑的奖励引导选择机制方面具有影响力,但主要的时间差异(TD)算法无法解释许多已在行为上证明的灵活或目标导向的行为。我们通过对比基于模型的RL算法来研究这种行为,因为它依赖于学习任务的地图或模型并在其中进行规划,而传统的无模型TD学习。为了区分人类的这些方法,我们在一个连续的空间导航任务中使用了功能磁共振成像,在这个任务中,迷宫布局的频繁变化迫使受试者不断地重新学习他们喜欢的路线,从而暴露了所采用的RL机制。我们通过比较选择行为和BOLD信号与从两种算法的模拟中提取的决策变量来寻找这种机制的神经基质的证据。纹状体中的选择和价值相关的BOLD信号,尽管最常与TD学习相关,但基于模型的理论可以更好地解释。此外,基于模型的值计算的前趋量与内侧颞叶和额叶皮层中的BOLD信号相关。这些结果表明,RL在大脑中的计算和解剖基底都有显著的扩展。
Although reinforcement learning (RL) theories have been influential in characterizing the brain’s mechanisms for reward-guided choice, the predominant temporal difference (TD) algorithm cannot explain many flexible or goal-directed actions that have been demonstrated behaviorally. We investigate such actions by contrasting an RL algorithm that is model-based, in that it relies on learning a map or model of the task and planning within it, to traditional model-free TD learning. To distinguish these approaches in humans, we used fMRI in a continuous spatial navigation task, in which frequent changes to the layout of the maze forced subjects continually to relearn their favored routes, thereby exposing the RL mechanisms employed. We sought evidence for the neural substrates of such mechanisms by comparing choice behavior and BOLD signals to decision variables extracted from simulations of either algorithm. Both choices and value-related BOLD signals in striatum, though most often associated with TD learning, were better explained by the model-based theory. Further, predecessor quantities for the model-based value computation were correlated with BOLD signals in the medial temporal lobe and frontal cortex. These results point to a significant extension of both the computational and anatomical substrates for RL in the brain.