Primate Orbitofrontal Cortex Codes Information Relevant for Managing Explore-Exploit Tradeoffs

Primate Orbitofrontal Cortex Codes Information Relevant for Managing Explore-Exploit Tradeoffs
复制标题

DOI:
10.1523/jneurosci.2355-19.2020
复制
发表时间:
2020-03-18
影响因子:
5.3
通讯作者:
Averbeck, Bruno B.
Averbeck, Bruno B.
中科院分区:
医学1区
文献类型:
--
作者:
Costa, Vincent D.;Averbeck, Bruno B.

文献摘要

被引文献

相似文献

强化学习是指为了获得奖赏和逃避惩罚而进行学习的行为过程。RL的一个重要组成部分是管理探索-利用权衡,它指的是在利用已知价值的选择和探索不熟悉的选择之间做出选择的问题。当三只雄性猴子执行三臂强盗学习任务时,我们在眼眶前额叶皮质(OFC)中检查了这种权衡的相关性,以及其他与RL相关的变量。在任务中,新的选择选项周期性地取代熟悉的选项。新选项的价值是未知的,猴子们不得不探索它们,看看它们是否比目前可用的其他选项更好。被选择的刺激和奖励结果的一致性被强烈编码在单个OFC神经元的反应中。这两个变量定义了我们模型中与决策相关的状态和状态转换。期权的选择价值和探索该期权的相对价值是在中间水平编码的。我们还发现,OFC值编码是刺激特定的,而不是与选项的身份无关的编码值。选择权的位置和当前环境的价值被以低水平编码。因此,我们发现了OFC中与学习和管理探索-利用权衡相关的变量的编码。这些结果与腹侧纹状体和杏仁核的研究结果一致,表明这个单突触连接的网络在基于选择的直接和未来后果的学习中发挥着重要作用。
Reinforcement learning (RL) refers to the behavioral process of learning to obtain reward and avoid punishment. An important component of RL is managing explore-exploit tradeoffs, which refers to the problem of choosing between exploiting options with known values and exploring unfamiliar options. We examined correlates of this tradeoff, as well as other RL related variables, in orbitofrontal cortex (OFC) while three male monkeys performed a three-armed bandit learning task. During the task, novel choice options periodically replaced familiar options. The values of the novel options were unknown, and the monkeys had to explore them to see if they were better than other currently available options. The identity of the chosen stimulus and the reward outcome were strongly encoded in the responses of single OFC neurons. These two variables define the states and state transitions in our model that are relevant to decision-making. The chosen value of the option and the relative value of exploring that option were encoded at intermediate levels. We also found that OFC value coding was stimulus specific, as opposed to coding value independent of the identity of the option. The location of the option and the value of the current environment were encoded at low levels. Therefore, we found encoding of the variables relevant to learning and managing explore-exploit tradeoffs in OFC. These results are consistent with findings in the ventral striatum and amygdala and show that this monosynaptically connected network plays an important role in learning based on the immediate and future consequences of choices.