Computational characteristics of the striatal dopamine system described by reinforcement learning with fast generalization

Computational characteristics of the striatal dopamine system described by reinforcement learning with fast generalization
复制标题

快速泛化强化学习描述纹状体多巴胺系统的计算特征

DOI:
10.1101/2019.12.12.873950
复制
发表时间:
2019
影响因子:
3.2
通讯作者:
Ishii Shin
Ishii Shin
中科院分区:
医学4区
文献类型:
--
作者:
Fujita Yoshihisa;Yagishita Sho;Kasai Haruo;Ishii Shin

文献摘要

相似文献

概括是将过去的经验应用于类似但不相同的情况的能力。它不仅影响刺激-结果关系,正如条件反射实验中所观察到的那样,而且可能对适应性行为至关重要,这涉及个人与环境之间的相互作用。计算建模可能会澄清泛化对自适应行为的影响,以及这种影响如何从底层计算中显现出来。最近的神经生物学观察表明,纹状体多巴胺系统实现泛化和随后的歧视,通过更新皮质纹状体突触连接的差异反应奖励和惩罚。在这项研究中,我们分析了这个神经生物学系统中的计算特性如何影响适应行为。提出了一种新的多层神经网络强化学习模型,该模型仅根据预测误差更新最后一层神经网络的突触权值。我们在输入层和隐藏层之间设置固定的连接,以保持隐藏层表示中输入的相似性。该网络实现了奖励和惩罚学习的快速泛化,从而促进了空间导航任务的安全和有效探索。值得注意的是,与不显示泛化的算法相比,它在早期学习阶段展示了快速的奖励方法和有效的惩罚规避。然而,干扰的网络,导致嘈杂的推广和受损的歧视诱导适应不良的估值。这些结果表明了纹状体多巴胺系统计算适应行为的优点和潜在的缺点。
Generalization is the ability to apply past experience to similar but non-identical situations. It not only affects stimulus-outcome relationships, as observed in conditioning experiments, but may also be essential for adaptive behaviors, which involve the interaction between individuals and their environment. Computational modeling could potentially clarify the effect of generalization on adaptive behaviors and how this effect emerges from the underlying computation. Recent neurobiological observation indicated that the striatal dopamine system achieves generalization and subsequent discrimination by updating the corticostriatal synaptic connections in differential response to reward and punishment. In this study, we analyzed how computational characteristics in this neurobiological system affects adaptive behaviors. We proposed a novel reinforcement learning model with multilayer neural networks in which the synaptic weights of only the last layer are updated according to the prediction error. We set fixed connections between the input and hidden layers to maintain the similarity of inputs in the hidden-layer representation. This network enabled fast generalization of reward and punishment learning, and thereby facilitated safe and efficient exploration of spatial navigation tasks. Notably, it demonstrated a quick reward approach and efficient punishment aversion in the early learning phase, compared to algorithms that do not show generalization. However, disturbance of the network that causes noisy generalization and impaired discrimination induced maladaptive valuation. These results suggested the advantage and potential drawback of computation by the striatal dopamine system with regard to adaptive behaviors.