Decomposing the effects of context valence and feedback information on speed and accuracy during reinforcement learning: a meta-analytical approach using diffusion decision modeling

Decomposing the effects of context valence and feedback information on speed and accuracy during reinforcement learning: a meta-analytical approach using diffusion decision modeling
复制标题

DOI:
10.3758/s13415-019-00723-1
复制
发表时间:
2019-06-01
影响因子:
2.9
通讯作者:
Lebreton, Mael
Lebreton, Mael
中科院分区:
医学3区
文献类型:
--
作者:
Fontanesi, Laura;Palminteri, Stefano;Lebreton, Mael

文献摘要

被引文献

相似文献

强化学习(RL)模型描述了人类和动物如何通过反复试验来选择最大限度地奖励和最大限度地减少惩罚的行动。传统的反应学习模型只关注选择,从而忽略了选择偏好和反应时间之间的交互作用,或者这些交互作用是如何受到语境因素的影响的。然而,在知觉决策领域,这种相互作用已被证明是区分不同潜在认知过程的重要因素。在这里,我们研究了这种相互作用,以揭示学习寻求回报和学习避免损失之间被忽视的差异。我们利用了来自四个RL实验的行为数据,这些实验的特点是操纵两个因素:结果效价(收益与损失)和反馈信息(部分与完全反馈)。贝叶斯元分析表明,这些语境因素对RTS和准确性的影响是不同的:价态只影响RTS,反馈信息同时影响RTS和准确性。为了分离潜在认知过程之间的关联,我们使用贝叶斯分层扩散决策模型(DDM)联合拟合了所有实验中的选择和RT。我们发现,反馈操作影响漂移率、阈值和非决定时间,这表明它不是单纯的难度效应。此外,效价影响非决定时间和阈值,表明在惩罚性情境中存在运动抑制。为了更好地了解学习动态,我们最终安装了RL和DDM的组合(RLDDM)。我们发现,虽然阈值受到特定于试验的决策冲突的影响,但非决策时间受到学习到的情境效价的影响。总体而言,我们的结果说明了在RL过程中联合建模RTS和选择数据的好处,以揭示不同学习环境下潜在决策的细微机制差异。
Reinforcement learning (RL) models describe how humans and animals learn by trial-and-error to select actions that maximize rewards and minimize punishments. Traditional RL models focus exclusively on choices, thereby ignoring the interactions between choice preference and response time (RT), or how these interactions are influenced by contextual factors. However, in the field of perceptual decision-making, such interactions have proven to be important to dissociate between different underlying cognitive processes. Here, we investigated such interactions to shed new light on overlooked differences between learning to seek rewards and learning to avoid losses. We leveraged behavioral data from four RL experiments, which feature manipulations of two factors: outcome valence (gains vs. losses) and feedback information (partial vs. complete feedback). A Bayesian meta-analysis revealed that these contextual factors differently affect RTs and accuracy: While valence only affects RTs, feedback information affects both RTs and accuracy. To dissociate between the latent cognitive processes, we jointly fitted choices and RTs across all experiments with a Bayesian, hierarchical diffusion decision model (DDM). We found that the feedback manipulation affected drift rate, threshold, and non-decision time, suggesting that it was not a mere difficulty effect. Moreover, valence affected non-decision time and threshold, suggesting a motor inhibition in punishing contexts. To better understand the learning dynamics, we finally fitted a combination of RL and DDM (RLDDM). We found that while the threshold was modulated by trial-specific decision conflict, the non-decision time was modulated by the learned context valence. Overall, our results illustrate the benefits of jointly modeling RTs and choice data during RL, to reveal subtle mechanistic differences underlying decisions in different learning contexts.