Multi-timescale reinforcement learning in the brain.

Multi-timescale reinforcement learning in the brain.
复制标题

大脑中的多时间尺度强化学习。

DOI:
10.1101/2023.11.12.566754
复制
发表时间:
2023
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
Uchida,Naoshige
Uchida,Naoshige
中科院分区:
--
文献类型:
--
作者:
Masset,Paul;Tano,Pablo;Kim,HyungGooR;Malik,AtharN;Pouget,Alexandre;Uchida,Naoshige

文献摘要

相似文献

为了在复杂的环境中茁壮成长,动物和人工智能必须学会适应性地采取行动,以最大限度地提高适应性和回报。这种自适应行为可以通过强化学习来学习1,这是一类成功训练人工代理2 -6和表征中脑多巴胺神经元放电的算法7 -9。在经典的强化学习中,智能体根据折扣因子控制的单个时间尺度以指数方式折扣未来的奖励。在这里,我们探讨了生物强化学习中多个时间尺度的存在。我们首先表明,强化代理在众多的时间尺度上学习具有明显的计算优势。接下来,我们报告说,多巴胺神经元在小鼠执行两个行为的任务编码奖励预测错误的折扣时间常数的多样性。我们的模型解释了线索诱发的瞬态反应和较慢的时间尺度波动(称为多巴胺斜坡)的时间折扣的异质性。至关重要的是,测量的单个神经元的折扣因子在两个任务中是相关的,这表明它是细胞特异性的。总之,我们的研究结果提供了一个新的范式来理解多巴胺神经元的功能异质性,这是人类和动物在许多情况下使用非指数折扣的经验观察的机制基础10 -14,并为设计更有效的强化学习算法开辟了新的途径。
To thrive in complex environments, animals and artificial agents must learn to act adaptively to maximize fitness and rewards. Such adaptive behavior can be learned through reinforcement learning1, a class of algorithms that has been successful at training artificial agents2–6 and at characterizing the firing of dopamine neurons in the midbrain7–9. In classical reinforcement learning, agents discount future rewards exponentially according to a single time scale, controlled by the discount factor. Here, we explore the presence of multiple timescales in biological reinforcement learning. We first show that reinforcement agents learning at a multitude of timescales possess distinct computational benefits. Next, we report that dopamine neurons in mice performing two behavioral tasks encode reward prediction error with a diversity of discount time constants. Our model explains the heterogeneity of temporal discounting in both cue-evoked transient responses and slower timescale fluctuations known as dopamine ramps. Crucially, the measured discount factor of individual neurons is correlated across the two tasks suggesting that it is a cell-specific property. Together, our results provide a new paradigm to understand functional heterogeneity in dopamine neurons, a mechanistic basis for the empirical observation that humans and animals use non-exponential discounts in many situations10–14, and open new avenues for the design of more efficient reinforcement learning algorithms.