Influences of Reinforcement and Choice Histories on Choice Behavior in Actor-Critic Learning

Influences of Reinforcement and Choice Histories on Choice Behavior in Actor-Critic Learning
复制标题

强化和选择历史对行动者批评学习中选择行为的影响

DOI:
10.1007/s42113-022-00145-2
复制
发表时间:
2022
期刊:
Computational Brain & Behavior
影响因子:
--
通讯作者:
Kimura Kenta
Kimura Kenta
中科院分区:
--
文献类型:
--
作者:
Katahira Kentaro;Kimura Kenta

文献摘要

参考文献

相似文献

强化学习模型已被用于神经科学和心理学领域的许多研究中,以模拟选择行为和潜在的计算过程。基于动作值的模型,表示动作的预期回报(例如,Q学习模型),通常用于此目的。同时,行动者-批评者学习模型,其中在单独的系统(分别为行动者和批评者)中执行给定状态的策略更新和预期奖励的评估,由于其能够解释生命系统的各种行为的特性而引起了关注。然而,模型行为的统计特性(即,选择如何取决于过去的奖励和选择)仍然难以捉摸。在这项研究中,我们研究的历史依赖性的演员-评论家模型的基础上,理论的考虑和数值模拟,同时考虑到与Q-学习模型的相似性和差异。我们发现,在行动者-批评者学习中,过去的奖励和选择之间的特定相互作用,这与Q学习不同,会影响当前的选择。我们还表明,行动者-批评者学习预测定性不同的行为从Q-学习,作为期望值越高,越不可能的行为将被选择之后。本研究通过阐明行动者-批评者学习在选择行为中的表现,为从行为中推断计算和心理学原理提供了有用的信息。
Reinforcement learning models have been used in many studies in the fields of neuroscience and psychology to model choice behavior and underlying computational processes. Models based on action values, which represent the expected reward from actions (e.g., Q-learning model), have been commonly used for this purpose. Meanwhile, the actor-critic learning model, in which the policy update and evaluation of an expected reward for a given state are performed in separate systems (actor and critic, respectively), has attracted attention due to its ability to explain the characteristics of various behaviors of living systems. However, the statistical property of the model behavior (i.e., how the choice depends on past rewards and choices) remains elusive. In this study, we examine the history dependence of the actor-critic model based on theoretical considerations and numerical simulations while considering the similarities with and differences from Q-learning models. We show that in actor-critic learning, a specific interaction between past reward and choice, which differs from Q-learning, influences the current choice. We also show that actor-critic learning predicts qualitatively different behavior from Q-learning, as the higher the expectation is, the less likely the behavior will be chosen afterwards. This study provides useful information for inferring computational and psychological principles from behavior by clarifying how actor-critic learning manifests in choice behavior.
DOI: 10.1111/pcn.13279
发表时间: 2021-09
影响因子: 11.9
作者:
Suzuki S;Yamashita Y;Katahira K
通讯作者: Katahira K
一种基于模型的选择偏好的简单计算算法
DOI: --
发表时间: 2017
期刊: Cognitive, Affective, & Behavioral Neuroscience
影响因子: --
作者:
Toyama;A.;Katahira;K.;& Ohira;H.
通讯作者: H.
DOI: 10.1016/j.neunet.2021.05.030
发表时间: 2021
期刊: Neural Networks
影响因子: 7.8
作者:
Ohta Hiroyuki;Satori Kuniaki;Takarada Yu;Arake Masashi;Ishizuka Toshiaki;Morimoto Yuji;Takahashi Tatsuji
通讯作者: Takahashi Tatsuji
DOI: 10.1371/journal.pone.0003795
发表时间: 2008
期刊: PloS one
影响因子: 3.7
作者:
Sakai Y;Fukai T
通讯作者: Fukai T