Actor–Critic Models of Reinforcement Learning in the Basal Ganglia: From Natural to Artificial Rats

Actor–Critic Models of Reinforcement Learning in the Basal Ganglia: From Natural to Artificial Rats
复制标题

基底神经节强化学习的演员-批评家模型:从自然大鼠到人工大鼠

DOI:
10.1177/105971230501300205
复制
发表时间:
2005
期刊:
影响因子:
1.6
通讯作者:
Agnès Guillot
Agnès Guillot
中科院分区:
计算机科学4区
文献类型:
--
作者:
M. Khamassi;Loïc Lachèze;Benoît Girard;A. Berthoz;Agnès Guillot

文献摘要

被引文献

相似文献

自1995年以来,许多强化学习的Actor-Critic架构已经被提出作为大鼠基底神经节中多巴胺样强化学习机制的模型。然而,这些模型通常在不同的任务中进行测试,因此很难比较它们对于自主动物的效率。我们在这里提出的比较四个架构的动画,因为它执行相同的奖励寻求任务。这将说明关于不同Actor子模块和Critic单元的管理的不同假设的结果,以及它们或多或少自主确定的协调。我们表明,经典的方法协调模块的混合专家,根据每个模块的性能,不允许解决我们的任务。然后,我们解决的问题,哪一个原则应有效地应用到联合收割机这些单位。最后以Psikharpax项目--一只需要在不可预测的环境中自主生存的人工鼠--为例,讨论了Actor-Critic模型在自然任务中的改进和准确性.
Since 1995, numerous Actor–Critic architectures for reinforcement learning have been proposed as models of dopamine-like reinforcement learning mechanisms in the rat’s basal ganglia. However, these models were usually tested in different tasks, and it is then difficult to compare their efficiency for an autonomous animat. We present here the comparison of four architectures in an animat as it per forms the same reward-seeking task. This will illustrate the consequences of different hypotheses about the management of different Actor sub-modules and Critic units, and their more or less autono mously determined coordination. We show that the classical method of coordination of modules by mixture of experts, depending on each module’s performance, did not allow solving our task. Then we address the question of which principle should be applied efficiently to combine these units. Improve ments for Critic modeling and accuracy of Actor–Critic models for a natural task are finally discussed in the perspective of our Psikharpax project—an artificial rat having to survive autonomously in unpre dictable environments.