Cortical delta activity reflects reward prediction error and related behavioral adjustments, but at different times

Cortical delta activity reflects reward prediction error and related behavioral adjustments, but at different times
复制标题

DOI:
10.1016/j.neuroimage.2015.02.007
复制
发表时间:
2015-04-15
期刊:
影响因子:
5.7
通讯作者:
Cavanagh, James F.
Cavanagh, James F.
中科院分区:
医学1区
文献类型:
--
作者:
Cavanagh, James F.

文献摘要

被引文献

相似文献

最近的研究表明,奖励预测错误会在头皮记录的脑电图(EEG)中引起正电压偏转;这种事件有时被称为奖励积极性。然而,对这一拟议关系的有力检验仍有待确定。其他重要的问题仍然没有得到解决:例如奖励积极性在预测未来最大化奖励的行为调整中的作用。为了回答这些问题,一个三臂强盗任务被用来调查的作用,积极的预测错误,在试验的探索和任务集为基础的开发。反馈锁定的奖励积极性的特点是δ带活动,这些相关的EEG功能与计算得出的积极预测误差的程度成比例。然而,这些现象也是分离的:计算模型预测剥削性的动作选择和相关的响应时间加快,而反馈锁定的EEG功能没有。令人信服的是,时间锁定到随后的强盗(P3)的δ带动力学成功地预测了这些行为。这些强盗锁定的研究结果包括增强顶叶运动皮层δ相位滞后,与反应时间加快的程度,这表明δ带活动在激励行动选择的机制作用。反馈与强盗锁定EEG信号的这种分离被解释为分层不同类型的预测误差的差异,在强化学习和决策过程中产生关于这些可分离的δ带现象的新预测。(C)2015 Elsevier Inc. All rights reserved.
Recent work has suggested that reward prediction errors elicit a positive voltage deflection in the scalp-recorded electroencephalogram (EEG); an event sometimes termed a reward positivity. However, a strong test of this proposed relationship remains to be defined. Other important questions remain unaddressed: such as the role of the reward positivity in predicting future behavioral adjustments that maximize reward. To answer these questions, a three-armed bandit task was used to investigate the role of positive prediction errors during trial-by-trial exploration and task-set based exploitation. The feedback-locked reward positivity was characterized by delta band activities, and these related EEG features scaled with the degree of a computationally derived positive prediction error. However, these phenomena were also dissociated: the computational model predicted exploitative action selection and related response time speeding whereas the feedback-locked EEG features did not. Compellingly, delta band dynamics time-locked to the subsequent bandit (the P3) successfully predicted these behaviors. These bandit-locked findings included an enhanced parietal to motor cortex delta phase lag that correlated with the degree of response time speeding, suggesting a mechanistic role for delta band activities in motivating action selection. This dissociation in feedback vs. bandit locked EEG signals is interpreted as a differentiation in hierarchically distinct types of prediction error, yielding novel predictions about these dissociable delta band phenomena during reinforcement learning and decision making. (C) 2015 Elsevier Inc. All rights reserved.