Contextual modulation of value signals in reward and punishment learning.

Contextual modulation of value signals in reward and punishment learning.
复制标题

DOI:
10.1038/ncomms9096
复制
发表时间:
2015-08-25
影响因子:
16.6
通讯作者:
Coricelli G
Coricelli G
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Palminteri S;Khamassi M;Joffily M;Coricelli G

文献摘要

被引文献

相似文献

与奖赏寻求相比,惩罚回避学习在计算和神经生物学水平上都不太清楚。在这里,我们证明,使用计算模型和功能磁共振成像在人类中,学习选项值在相对的上下文依赖规模提供了一个简单的计算解决方案,避免学习。上下文(或状态)值设置了一个参考点,在更新选项值之前,应该将结果与该参考点进行比较。因此,在总体预期值为负的背景下,成功的惩罚回避获得了正的价值,从而加强了反应。学习后的选项价值评估显示,当被试被告知放弃的替代品的结果(反事实信息)时,情境影响得到增强。这反映在神经水平上的负面结果编码从前额叶到腹侧纹状体的转变,表明价值情境化也限制了调动对手惩罚学习系统的需要。 与学习理论的预测相反,人类学会了同样好地寻求奖励和避免惩罚。在这里,作者提供了一个优雅的解决方案,通过证明人类学习选项值相对于一个参考点,由一个共同的神经基板。
Compared with reward seeking, punishment avoidance learning is less clearly understood at both the computational and neurobiological levels. Here we demonstrate, using computational modelling and fMRI in humans, that learning option values in a relative—context-dependent—scale offers a simple computational solution for avoidance learning. The context (or state) value sets the reference point to which an outcome should be compared before updating the option value. Consequently, in contexts with an overall negative expected value, successful punishment avoidance acquires a positive value, thus reinforcing the response. As revealed by post-learning assessment of options values, contextual influences are enhanced when subjects are informed about the result of the forgone alternative (counterfactual information). This is mirrored at the neural level by a shift in negative outcome encoding from the anterior insula to the ventral striatum, suggesting that value contextualization also limits the need to mobilize an opponent punishment learning system. In contrast to predictions from learning theory, humans learn to seek rewards and avoid punishments equally well. Here the authors offer an elegant solution to this problem by demonstrating that humans learn option values relative to a reference point subserved by a common neural substrate.