Enhancing reinforcement learning models by including direct and indirect pathways improves performance on striatal dependent tasks.

Enhancing reinforcement learning models by including direct and indirect pathways improves performance on striatal dependent tasks.
复制标题

通过包含直接和间接路径来增强强化学习模型可以提高纹状体依赖任务的绩效。

DOI:
10.1371/journal.pcbi.1011385
复制
发表时间:
2023-08
影响因子:
4.3
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

在理解学习行为方面的一项重大进展源于一些实验,这些实验表明奖赏学习需要多巴胺输入到纹状体神经元,并且源于皮质 - 纹状体突触的突触可塑性。众多强化学习模型通过利用奖赏预测误差来模拟这种依赖多巴胺的突触可塑性,奖赏预测误差类似于多巴胺神经元的放电,以此来学习针对一组线索的最佳行动。尽管这些模型能够解释行为的许多方面,但要重现某些类型的目标导向行为,如恢复和逆转,则需要额外的模型组件。在此,我们提出一种强化学习模型TD2Q,它通过两个Q矩阵与基底神经节有更好的对应关系,其中一个矩阵代表直接通路神经元(G),另一个代表间接通路神经元(N)。与以往的双Q架构不同,TD2Q一个新颖且关键的方面是利用时间差分奖赏预测误差来更新G矩阵和N矩阵。使用带有依赖奖赏的自适应探索参数的softmax函数为N和G选择最佳行动,然后通过应用于两种行动概率的第二个选择步骤来解决差异。该模型在一系列多步骤任务中进行了测试,包括消退、恢复、辨别;切换奖赏概率学习;以及序列学习。模拟结果表明,TD2Q在选择和序列学习任务中产生的行为与啮齿动物类似,并且学习多步骤任务需要使用时间差分奖赏预测误差。正如实验中所观察到的,阻断N矩阵的更新规则会阻碍辨别学习。使用两个矩阵时,序列学习任务的表现得到显著提升。这些结果表明,纳入基底神经节生理学的更多方面能够提高强化学习模型的性能,更好地重现动物行为,并为了解直接和间接通路纹状体神经元的作用提供见解。 人类和动物在仅以正确行为的奖赏作为唯一反馈时,极其擅长学习执行复杂任务。学习的早期阶段以探索可能的行动为特征,而后期阶段则以优化行动序列为特征。实验证据表明,奖赏由多巴胺信号编码,并且多巴胺也会影响探索程度。强化学习算法是一类机器学习算法,它们利用奖赏信号来确定采取行动的价值。这些算法与基底神经节的信息处理有一定相似性,并且能够解释几种类型的学习行为。我们对其中一种算法——Q学习进行扩展,以增强其与基底神经节回路的相似性,并在多个学习任务中评估其性能。我们表明,通过纳入两条相反的基底神经节通路,我们能够提高在操作性条件反射任务和一项困难的序列学习任务中的表现。这些结果表明,纳入大脑回路的更多方面可能会进一步提升强化学习算法的性能。
A major advance in understanding learning behavior stems from experiments showing that reward learning requires dopamine inputs to striatal neurons and arises from synaptic plasticity of cortico-striatal synapses. Numerous reinforcement learning models mimic this dopamine-dependent synaptic plasticity by using the reward prediction error, which resembles dopamine neuron firing, to learn the best action in response to a set of cues. Though these models can explain many facets of behavior, reproducing some types of goal-directed behavior, such as renewal and reversal, require additional model components. Here we present a reinforcement learning model, TD2Q, which better corresponds to the basal ganglia with two Q matrices, one representing direct pathway neurons (G) and another representing indirect pathway neurons (N). Unlike previous two-Q architectures, a novel and critical aspect of TD2Q is to update the G and N matrices utilizing the temporal difference reward prediction error. A best action is selected for N and G using a softmax with a reward-dependent adaptive exploration parameter, and then differences are resolved using a second selection step applied to the two action probabilities. The model is tested on a range of multi-step tasks including extinction, renewal, discrimination; switching reward probability learning; and sequence learning. Simulations show that TD2Q produces behaviors similar to rodents in choice and sequence learning tasks, and that use of the temporal difference reward prediction error is required to learn multi-step tasks. Blocking the update rule on the N matrix blocks discrimination learning, as observed experimentally. Performance in the sequence learning task is dramatically improved with two matrices. These results suggest that including additional aspects of basal ganglia physiology can improve the performance of reinforcement learning models, better reproduce animal behaviors, and provide insight as to the role of direct- and indirect-pathway striatal neurons. Humans and animals are exceedingly adept at learning to perform complicated tasks when the only feedback is reward for correct actions. Early phases of learning are characterized by exploration of possible actions, and later phases of learning are characterized by optimizing the action sequence. Experimental evidence suggests that reward is encoded by the dopamine signal, and that dopamine also can influence the degree of exploration. Reinforcement learning algorithms are machine learning algorithms that use the reward signal to determine the value of taking an action. These algorithms have some similarity to information processing by the basal ganglia, and can explain several types of learning behavior. We extend one of these algorithms, Q learning, to increase the similarity to basal ganglia circuitry, and evaluate performance on several learning tasks. We show that by incorporating two opposing basal ganglia pathways, we can improve performance on operant conditioning tasks and a difficult sequence learning task. These results suggest that incorporating additional aspects of brain circuitry could further improve performance of reinforcement learning algorithms.
DOI: 10.1016/j.neuroscience.2009.03.015
发表时间: 2009-06-02
期刊: NEUROSCIENCE
影响因子: 3.3
作者:
Fino, E.;Paille, V.;Venance, L.
通讯作者: Venance, L.
DOI: 10.1037/a0037015
发表时间: 2014-07-01
影响因子: 5.4
作者:
Collins, Anne G. E.;Frank, Michael J.
通讯作者: Frank, Michael J.
DOI: 10.1073/pnas.1613337113
发表时间: 2016-10-04
影响因子: 11.1
作者:
Crittenden, Jill R.;Tillberg, Paul W.;Graybiel, Ann M.
通讯作者: Graybiel, Ann M.
DOI: 10.3389/fncel.2021.639082
发表时间: 2021
影响因子: 5.3
作者:
Gorodetski L;Loewenstern Y;Faynveitz A;Bar-Gad I;Blackwell KT;Korngreen A
通讯作者: Korngreen A
DOI: 10.1098/rstb.2007.2098
发表时间: 2007-05-29
影响因子: 6.3
作者:
Cohen, Jonathan D.;McClure, Samuel M.;Yu, Angela J.
通讯作者: Yu, Angela J.