Model-based learning retrospectively updates model-free values.

Model-based learning retrospectively updates model-free values.
复制标题

DOI:
10.1038/s41598-022-05567-3
复制
发表时间:
2022-02-11
期刊:
影响因子:
4.6
通讯作者:
Manohar SG
Manohar SG
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Doody M;Van Swieten MMH;Manohar SG

文献摘要

参考文献

相似文献

强化学习(RL)被广泛认为可以分为两种不同的计算策略。无模型学习是一个简单的RL过程,其中值与动作相关联,而基于模型的学习依赖于环境内部模型的形成,以最大化奖励。最近,理论和动物研究表明,这种模型可以用来训练无模型行为,减少昂贵的前瞻性规划的负担。在这里,我们设计了一种方法来探索人类行为中的这种可能性。我们采用了一个两阶段的决策任务,并发现证据表明,在学习时基于模型的过程可以改变健康个体的无模型估值。我们要求人们对一个不相关的特征的主观价值进行评级,这个特征是在做出基于模型的决定时看到的。这些不相关的特征值评级通过奖励进行更新,但在某种程度上,考虑到所选择的行动是否应该被回顾性地采取。这种基于模型的对无模型价值评级的影响最好通过奖励预测误差来解释,该误差是相对于最有可能导致奖励的决策路径计算的。这种效应独立于注意力而发生,当参与者没有被明确告知环境的结构时,这种效应就不存在。这些研究结果表明,目前的概念模型为基础的和无模型的学习需要更新,有利于一个更综合的方法。本研究为今后进一步研究这两种学习系统之间的对话提供了一个实证性的工具。
Reinforcement learning (RL) is widely regarded as divisible into two distinct computational strategies. Model-free learning is a simple RL process in which a value is associated with actions, whereas model-based learning relies on the formation of internal models of the environment to maximise reward. Recently, theoretical and animal work has suggested that such models might be used to train model-free behaviour, reducing the burden of costly forward planning. Here we devised a way to probe this possibility in human behaviour. We adapted a two-stage decision task and found evidence that model-based processes at the time of learning can alter model-free valuation in healthy individuals. We asked people to rate subjective value of an irrelevant feature that was seen at the time a model-based decision would have been made. These irrelevant feature value ratings were updated by rewards, but in a way that accounted for whether the selected action retrospectively ought to have been taken. This model-based influence on model-free value ratings was best accounted for by a reward prediction error that was calculated relative to the decision path that would most likely have led to the reward. This effect occurred independently of attention and was not present when participants were not explicitly told about the structure of the environment. These findings suggest that current conceptions of model-based and model-free learning require updating in favour of a more integrated approach. Our task provides an empirical handle for further study of the dialogue between these two learning systems in the future.
DOI: 10.1016/j.biopsych.2018.12.017
发表时间: 2019-06-01
影响因子: 10.6
作者:
Groman, Stephanie M.;Massi, Bart;Taylor, Jane R.
通讯作者: Taylor, Jane R.
基于模型和无模型学习之间仲裁的基础神经计算。
DOI: 10.1016/j.neuron.2013.11.028
发表时间: 2014-02-05
期刊: Neuron
影响因子: 16.2
作者:
Lee SW;Shimojo S;O'Doherty JP
通讯作者: O'Doherty JP
DOI: 10.1126/science.abf1357
发表时间: 2021-05-21
期刊: Science (New York, N.Y.)
影响因子: --
作者:
Liu Y;Mattar MG;Behrens TEJ;Daw ND;Dolan RJ
通讯作者: Dolan RJ
DOI: 10.1016/j.neuron.2011.02.027
发表时间: 2011-03-24
期刊: Neuron
影响因子: 16.2
作者:
Daw ND;Gershman SJ;Seymour B;Dayan P;Dolan RJ
通讯作者: Dolan RJ
DOI: 10.3758/s13415-016-0487-3
发表时间: 2017-04-01
影响因子: 2.9
作者:
Eppinger, Ben;Walter, Maik;Li, Shu-Chen
通讯作者: Li, Shu-Chen