Temporal dynamics of prediction error processing during reward-based decision making

Temporal dynamics of prediction error processing during reward-based decision making
复制标题

DOI:
10.1016/j.neuroimage.2010.05.052
复制
发表时间:
2010-10-15
期刊:
影响因子:
5.7
通讯作者:
Heekeren, Hauke R.
Heekeren, Hauke R.
中科院分区:
医学1区
文献类型:
--
作者:
Philiastides, Marios G.;Biele, Guido;Heekeren, Hauke R.

文献摘要

被引文献

相似文献

适应性决策依赖于与潜在选择相关的奖励的准确表示。这些表征可以通过强化学习(RL)机制获得,RL机制使用预测误差(PE,期望和收到的奖励之间的差异)作为学习信号来更新奖励预期。虽然脑电图实验强调了反馈相关电位在行为监测中的作用,但关于反馈处理的时间序列和反馈相关电位在基于奖励的决策过程中的具体功能的重要问题仍然存在。在这里,我们假设反馈处理从结果效价的定性评估开始,随后由PE大小的定量表示补充。一项基于模型的单次试验分析结果显示,在反向学习任务期间收集的脑电图数据显示,反馈结果在大约220毫秒后,就其效价(正与负)进行了初步分类评估。在300毫秒左右,与维持的价值评估并行,大脑也表示关于PE大小的定量信息,从而提供更新奖励预期和指导适应性决策所需的完整信息。重要的是,我们基于RL模型的pe的单次脑电图分析表明,反馈相关电位不仅反映了错误意识,而且反映了对学习奖励偶然性至关重要的定量信息。(C) 2010爱思唯尔公司版权所有。
Adaptive decision making depends on the accurate representation of rewards associated with potential choices. These representations can be acquired with reinforcement learning (RL) mechanisms, which use the prediction error (PE, the difference between expected and received rewards) as a learning signal to update reward expectations. While EEG experiments have highlighted the role of feedback-related potentials during performance monitoring, important questions about the temporal sequence of feedback processing and the specific function of feedback-related potentials during reward-based decision making remain. Here, we hypothesized that feedback processing starts with a qualitative evaluation of outcome-valence, which is subsequently complemented by a quantitative representation of PE magnitude. Results of a model-based single-trial analysis of EEG data collected during a reversal learning task showed that around 220 ms after feedback outcomes are initially evaluated categorically with respect to their valence (positive vs. negative). Around 300 ms, and parallel to the maintained valence-evaluation, the brain also represents quantitative information about PE magnitude, thus providing the complete information needed to update reward expectations and to guide adaptive decision making. Importantly, our single-trial EEG analysis based on PEs from an RL model showed that the feedback-related potentials do not merely reflect error awareness, but rather quantitative information crucial for learning reward contingencies. (C) 2010 Elsevier Inc. All rights reserved.