Signals in human striatum are appropriate for policy update rather than value prediction.

Signals in human striatum are appropriate for policy update rather than value prediction.
复制标题

DOI:
10.1523/jneurosci.6316-10.2011
复制
发表时间:
2011-04-06
期刊:
The Journal of neuroscience : the official journal of the Society for Neuroscience
影响因子:
--
通讯作者:
Daw ND
Daw ND
中科院分区:
其他
文献类型:
--
作者:
Li J;Daw ND

文献摘要

被引文献

相似文献

有影响力的强化学习(RL)理论提出,大脑黑质纹状体系统中的“预测错误”信号指导学习试错决策。然而,由于不同的决策变量可以从数量上相似的误差信号中学习,一个关键的问题是由误差信号训练的决策表示的内容是什么。我们使用功能性磁共振成像(fMRI)来监测双臂强盗反事实决策任务中的神经活动,该任务为人类受试者提供有关先前的信息以及获得的货币结果,以便分离更新每个动作的预期值的教学信号,与训练动作之间的相对偏好(“政策”)的信号。两种选择的奖励概率彼此独立地变化。这种特定的设计使我们能够测试受试者的选择行为是否受到基于策略的方法的指导,这些方法直接将状态映射到有利的行动,或者基于价值的方法,如Q学习,其中选择策略是通过学习中间表示(奖励期望)来生成的。在行为上,我们发现人类参与者的选择受到先前试验中获得的和放弃的奖励的显著影响。我们还发现,受试者纹状体的血氧水平依赖性(BOLD)反应受到经验奖励和放弃奖励的反向调制,而不受奖励期望的调制。这种神经模式,以及受试者的选择行为,是一致的教学信号发展的“习惯”或相对的行动偏好,而不是预测错误更新单独的行动值。
Influential reinforcement learning (RL) theories propose that “prediction error” signals in the brain’s nigrostriatal system guide learning for trial-and-error decision-making. However, since different decision variables can be learned from quantitatively similar error signals, a critical question is what is the content of decision representations trained by the error signals. We used functional magnetic resonance imaging (fMRI) to monitor neural activity in a two-armed-bandit counterfactual decision task that provided human subjects with information about foregone as well as obtained monetary outcomes so as to dissociate teaching signals that update expected values for each action, vs. signals that train relative preferences between actions (a “policy”). The reward probabilities of both choices varied independently from each other. This specific design allowed us to test whether subjects’ choice behavior was guided by policy-based methods, which directly map states to advantageous actions, or value-based methods such as Q-learning, where choice policies are instead generated by learning an intermediate representation (reward expectancy). Behaviorally, we found human participants’ choices were significantly influenced by obtained as well as foregone rewards from the previous trial. We also found subjects’ blood-oxygen-level-dependent (BOLD) responses in striatum were modulated in opposite directions by the experienced and foregone rewards but not by reward expectancy. This neural pattern, as well as subjects’ choice behavior, is consistent with a teaching signal for developing “habits” or relative action preferences, rather than prediction errors for updating separate action values.