Hippocampal replay contributes to within session learning in a temporal difference reinforcement learning model

Hippocampal replay contributes to within session learning in a temporal difference reinforcement learning model
复制标题

DOI:
10.1016/j.neunet.2005.08.009
复制
发表时间:
2005-11-01
期刊:
影响因子:
7.8
通讯作者:
Redish, AD
Redish, AD
中科院分区:
计算机科学1区
文献类型:
--
作者:
Johnson, A;Redish, AD

文献摘要

被引文献

相似文献

时间差分强化学习(TDRL)算法被假设为部分解释基底节功能,但学习速度比真实动物慢。改进的TDRL算法(例如DYNA-Q系列)通过离线练习经验序列,比标准TDRL学习更快。我们认为,重放现象,即海马神经元集合在随后的休息和睡眠中重播先前经历的放电序列,可能提供练习序列来提高TDRL学习的速度,即使在一个单独的会话中也是如此。我们在一个多T选择任务的计算模型中检验了这一假设的合理性。老鼠在这项任务中表现出两种学习速度:错误的快速减少和刻板印象的缓慢发展。将开发重放添加到模型可以加速学习正确的路径,但会减慢对该路径的刻板印象。这些模型提供了关于海马体失活以及海马体重放对这一任务的影响的可测试预测。(C)2005爱思唯尔有限公司。保留所有权利。
Temporal difference reinforcement learning (TDRL) algorithms, hypothesized to partially explain basal ganglia functionality, learn more slowly than real animals. Modified TDRL algorithms (e.g. the Dyna-Q family) learn faster than standard TDRL by practicing experienced sequences offline. We suggest that the replay phenomenon, in which ensembles of hippocampal neurons replay previously experienced firing sequences during subsequent rest and sleep, may provide practice sequences to improve the speed of TDRL learning, even within a single session. We test the plausibility of this hypothesis in a computational model of a multiple-T choice-task. Rats show two learning rates on this task: a fast decrease in errors and a slow development of a stereotyped path. Adding developing replay to the model accelerates learning the correct path, but slows down the stereotyping of that path. These models provide testable predictions relating the effects of hippocampal inactivation as well as hippocampal replay on this task. (c) 2005 Elsevier Ltd. All rights reserved.