Whole-Brain Neural Dynamics of Probabilistic Reward Prediction.

Whole-Brain Neural Dynamics of Probabilistic Reward Prediction.
复制标题

DOI:
10.1523/jneurosci.2943-16.2017
复制
发表时间:
2017-04-05
期刊:
The Journal of neuroscience : the official journal of the Society for Neuroscience
影响因子:
--
通讯作者:
Dolan RJ
Dolan RJ
中科院分区:
其他
文献类型:
--
作者:
Bach DR;Symmonds M;Barnes G;Dolan RJ

文献摘要

被引文献

相似文献

预测未来的奖励对于执行最佳行动至关重要。虽然我们知道大脑中有许多区域对这种预测进行编码,但对相关表征如何随时间演变的详细描述却缺乏。在这里,我们使用人类脑磁图(MEG)和重建源的瞬时活动的多变量分析来解决这个问题。我们在一个简单的工具奖励学习任务上对参与者进行了过度训练,在这个任务中,几何线索预测了可能的奖励分布,并在2000毫秒后从中揭示了样本。我们表明,预测的平均回报(即期望值)和预测的回报变异性(即经济风险)是明显编码的。早期,平均奖励的表征出现在顶叶和视觉区,随后出现在额叶区,眶额叶皮层最后出现。值得注意的是,奖励变异性编码同时出现在顶叶/感觉和额叶来源,并且晚于平均奖励编码。眼窝额叶可变性编码与平均奖励编码同时出现。至关重要的是,交叉预测表明平均奖励和可变性表征是不同的,同时也表明瞬时表征随着时间的推移变得更加稳定。在各种来源中,可变性信号的最佳拟合指标是变异系数(而不是SD或方差),但对于单个大脑区域,可以看到不同的最佳指标。我们的数据展示了概率奖励预测的动态编码是如何在时间和空间上在大脑中展开的。预测未来的奖励对最佳行为至关重要。为了深入了解潜在的神经计算,我们研究了大脑中的奖励表征是如何随着时间的推移而产生的。使用脑磁图,我们发现预测的平均奖励的表现出现在顶叶/感觉区域较早,较晚在额叶皮层。相比之下,预测奖励可变性表征在大多数区域同时出现,比平均奖励稍晚。对于这两个特性,表示在稳定之前会动态变化1000毫秒。编码可变性的最佳度量是变异系数,在大脑区域之间可以看到这种编码的异质性。研究结果为预测性奖励表征的出现提供了新的见解。
Predicting future reward is paramount to performing an optimal action. Although a number of brain areas are known to encode such predictions, a detailed account of how the associated representations evolve over time is lacking. Here, we address this question using human magnetoencephalography (MEG) and multivariate analyses of instantaneous activity in reconstructed sources. We overtrained participants on a simple instrumental reward learning task where geometric cues predicted a distribution of possible rewards, from which a sample was revealed 2000 ms later. We show that predicted mean reward (i.e., expected value), and predicted reward variability (i.e., economic risk), are encoded distinctly. Early on, representations of mean reward are seen in parietal and visual areas, and later in frontal regions with orbitofrontal cortex emerging last. Strikingly, an encoding of reward variability emerges simultaneously in parietal/sensory and frontal sources and later than mean reward encoding. An orbitofrontal variability encoding emerged around the same time as that seen for mean reward. Crucially, cross-prediction showed that mean reward and variability representations are distinct and also revealed that instantaneous representations become more stable over time. Across sources, the best fitting metric for variability signals was coefficient of variation (rather than SD or variance), but distinct best metrics were seen for individual brain regions. Our data demonstrate how a dynamic encoding of probabilistic reward prediction unfolds in the brain both in time and space. SIGNIFICANCE STATEMENT Predicting future reward is paramount to optimal behavior. To gain insight into the underlying neural computations, we investigate how reward representations in the brain arise over time. Using magnetoencephalography, we show that a representation of predicted mean reward emerges early in parietal/sensory regions and later in frontal cortex. In contrast, predicted reward variability representations appear in most regions at the same time, and slightly later than for mean reward. For both features, representations dynamically change >1000 ms before stabilizing. The best metric for encoding variability is coefficient of variation, with heterogeneity in this encoding seen between brain areas. The results provide novel insights into the emergence of predictive reward representations.