Reward-based training of recurrent neural networks for cognitive and value-based tasks

Reward-based training of recurrent neural networks for cognitive and value-based tasks
复制标题

DOI:
10.7554/elife.21492
复制
发表时间:
2016-08
期刊:
影响因子:
7.7
通讯作者:
H. F. Song;G. R. Yang;Xiao-Jing Wang;Xiao-Jing Wang
H. F. Song;G. R. Yang;Xiao-Jing Wang;Xiao-Jing Wang
中科院分区:
生物学1区
文献类型:
--
作者:
H. F. Song;G. R. Yang;Xiao-Jing Wang;Xiao-Jing Wang

文献摘要

被引文献

相似文献

经过训练的神经网络模型表现出在行为动物的神经记录中观察到的许多特征,并且其活动和连接可以被充分分析,可以提供对神经机制的见解。然而,与常用的从分级错误信号进行监督学习的方法相反,动物通过强化学习从对确定行为的奖励反馈中学习。当最佳行为取决于动物对信心或主观偏好的内部判断时,奖励最大化特别相关。在这里,我们描述了递归神经网络的基于奖励的训练,其中价值网络通过使用策略网络的选定动作和活动来预测未来的奖励来指导学习。我们发现,这些模型捕获的行为和电生理研究结果,从众所周知的实验范例。我们的研究结果提供了一个统一的框架,调查不同的认知和基于价值的计算,包括价值表征的作用,是必不可少的学习,但不执行任务。
Trained neural network models, which exhibit many features observed in neural recordings from behaving animals and whose activity and connectivity can be fully analyzed, may provide insights into neural mechanisms. In contrast to commonly used methods for supervised learning from graded error signals, however, animals learn from reward feedback on definite actions through reinforcement learning. Reward maximization is particularly relevant when the optimal behavior depends on an animal’s internal judgment of confidence or subjective preferences. Here, we describe reward-based training of recurrent neural networks in which a value network guides learning by using the selected actions and activity of the policy network to predict future reward. We show that such models capture both behavioral and electrophysiological findings from well-known experimental paradigms. Our results provide a unified framework for investigating diverse cognitive and value-based computations, including a role for value representation that is essential for learning, but not executing, a task.