How much of reinforcement learning is working memory, not reinforcement learning? A behavioral, computational, and neurogenetic analysis.

How much of reinforcement learning is working memory, not reinforcement learning? A behavioral, computational, and neurogenetic analysis.
复制标题

DOI:
10.1111/j.1460-9568.2011.07980.x
复制
发表时间:
2012-04
期刊:
The European journal of neuroscience
影响因子:
--
通讯作者:
Frank MJ
Frank MJ
中科院分区:
其他
文献类型:
--
作者:
Collins AG;Frank MJ

文献摘要

参考文献

被引文献

相似文献

工具性学习涉及皮质纹状体回路和多巴胺能系统。该系统通常在强化学习(RL)框架中通过增量累积状态和动作的奖励值来建模。然而,人类的学习也涉及前额叶皮层机制参与高层次的认知功能。这些系统的相互作用仍然知之甚少,人类行为模型经常忽略工作记忆(WM),因此错误地将行为差异分配给RL系统。在这里,我们设计了一个任务,突出了这两个过程的深刻纠缠,即使是在简单的学习问题中。通过系统地改变学习问题的大小和刺激重复之间的延迟,我们分别提取了负载和延迟对学习的wm特异性影响。我们提出了一个新的计算模型来解释在被试行为中观察到的RL和WM过程的动态整合。将有能力限制的WM整合到模型中,使我们能够捕获在纯RL框架中无法捕获的行为差异,即使我们(令人难以置信地)允许为每个集合大小单独的RL系统。WM组件还允许对单个RL过程进行更合理的估计。最后,我们报告了对前额叶和基底神经节功能具有相对特异性的两种遗传多态性的影响。编码儿茶酚- o -甲基转移酶的COMT基因选择性地影响WM能力的模型估计,而编码g蛋白偶联受体6的GPR6基因影响RL学习率。因此,这项研究使我们能够明确高水平和低水平认知功能对工具性学习的不同影响,而不仅仅是简单的强化学习模型所提供的可能性。
Instrumental learning involves corticostriatal circuitry and the dopaminergic system. This system is typically modeled in the reinforcement learning (RL) framework by incrementally accumulating reward values of states and actions. However, human learning also implicates prefrontal cortical mechanisms involved in higher level cognitive functions. The interaction of these systems remains poorly understood, and models of human behavior often ignore working memory (WM) and therefore incorrectly assign behavioral variance to the RL system. Here we designed a task that highlights the profound entanglement of these two processes, even in simple learning problems. By systematically varying the size of the learning problem and delay between stimulus repetitions, we separately extracted WM-specific effects of load and delay on learning. We propose a new computational model that accounts for the dynamic integration of RL and WM processes observed in subjects' behavior. Incorporating capacity-limited WM into the model allowed us to capture behavioral variance that could not be captured in a pure RL framework even if we (implausibly) allowed separate RL systems for each set size. The WM component also allowed for a more reasonable estimation of a single RL process. Finally, we report effects of two genetic polymorphisms having relative specificity for prefrontal and basal ganglia functions. Whereas the COMT gene coding for catechol-O-methyl transferase selectively influenced model estimates of WM capacity, the GPR6 gene coding for G-protein-coupled receptor 6 influenced the RL learning rate. Thus, this study allowed us to specify distinct influences of the high-level and low-level cognitive functions on instrumental learning, beyond the possibilities offered by simple RL models.
DOI: 10.1073/pnas.111134598
发表时间: 2001-06-05
影响因子: 11.1
作者:
Egan, MF;Goldberg, TE;Weinberger, DR
通讯作者: Weinberger, DR
DOI: 10.1177/0963721409359277
发表时间: 2010-02-01
影响因子: 7.2
作者:
Cowan N
通讯作者: Cowan N
DOI: 10.1162/jocn.2009.21318
发表时间: 2010-07-01
影响因子: 3.2
作者:
de Frias, Cindy M.;Marklund, Petter;Nyberg, Lars
通讯作者: Nyberg, Lars
DOI: 10.1016/j.neuron.2011.02.027
发表时间: 2011-03-24
期刊: Neuron
影响因子: 16.2
作者:
Daw ND;Gershman SJ;Seymour B;Dayan P;Dolan RJ
通讯作者: Dolan RJ
DOI: 10.1007/s10048-007-0084-2
发表时间: 2007-08-01
期刊: NEUROGENETICS
影响因子: 2.2
作者:
Ernst, Carl;Sequeira, Adolfo;Turecki, Gustavo
通讯作者: Turecki, Gustavo