The asymmetric learning rates of murine exploratory behavior in sparse reward environments

The asymmetric learning rates of murine exploratory behavior in sparse reward environments
复制标题

稀疏奖励环境中小鼠探索行为的不对称学习率

DOI:
10.1016/j.neunet.2021.05.030
复制
发表时间:
2021
期刊:
影响因子:
7.8
通讯作者:
Takahashi Tatsuji
Takahashi Tatsuji
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ohta Hiroyuki;Satori Kuniaki;Takarada Yu;Arake Masashi;Ishizuka Toshiaki;Morimoto Yuji;Takahashi Tatsuji

文献摘要

被引文献

相似文献

动物的目标导向行为可以通过强化学习算法来建模。这种算法利用动作值来预测所选动作的未来结果,并根据积极和消极的结果更新这些值。在动物行为的许多模型中,动作值是基于一个共同的学习率对称地更新的,也就是说,对于积极和消极的结果,以同样的方式更新。然而,在缺乏奖励的环境中,动物的学习率可能不均衡。为了研究奖励和非奖励中学习率的不对称性,我们使用具有不同学习率的q -学习模型分析了小鼠在五臂强盗任务中的探索行为。稀缺奖励环境下的积极学习率显著高于丰富奖励环境下的积极学习率,反之,稀缺环境下的消极学习率显著低于丰富奖励环境下的消极学习率。在稀缺环境下,正负学习率比约为10,在丰富环境下,正负学习率比约为2。这一结果表明,当奖励概率较低时,小鼠倾向于忽略失败,并利用罕见的奖励。计算模型分析表明,学习率的增加可能导致对稀有奖励事件的高估和坚持,增加了稀缺环境下的总奖励获取,但不利于公正的探索。
Goal-oriented behaviors of animals can be modeled by reinforcement learning algorithms. Such algorithms predict future outcomes of selected actions utilizing action values and updating those values in response to the positive and negative outcomes. In many models of animal behavior, the action values are updated symmetrically based on a common learning rate, that is, in the same way for both positive and negative outcomes. However, animals in environments with scarce rewards may have uneven learning rates. To investigate the asymmetry in learning rates in reward and non-reward, we analyzed the exploration behavior of mice in five-armed bandit tasks using a Q-learning model with differential learning rates for positive and negative outcomes. The positive learning rate was significantly higher in a scarce reward environment than in a rich reward environment, and conversely, the negative learning rate was significantly lower in the scarce environment. The positive to negative learning rate ratio was about 10 in the scarce environment and about 2 in the rich environment. This result suggests that when the reward probability was low, the mice tend to ignore failures and exploit the rare rewards. Computational modeling analysis revealed that the increased learning rates ratio could cause an overestimation of and perseveration on rare-rewarding events, increasing total reward acquisition in the scarce environment but disadvantaging impartial exploration.