Speed/accuracy trade-off between the habitual and the goal-directed processes.

Speed/accuracy trade-off between the habitual and the goal-directed processes.
复制标题

DOI:
10.1371/journal.pcbi.1002055
复制
发表时间:
2011-05
影响因子:
4.3
通讯作者:
Piray P
Piray P
中科院分区:
生物学2区
文献类型:
--
作者:
Keramati M;Dezfouli A;Piray P

文献摘要

参考文献

被引文献

相似文献

工具性反应被假设为两种类型:习惯性反应和目标导向反应,分别由感觉运动回路和联合皮质-基底节回路介导。可以假设,这两种不同的联想学习机制的存在源于它们在不同学习阶段所具有的比较优势。在本文中,我们假设目标导向系统在行为上是灵活的,但在选择上是缓慢的。相比之下,习惯性系统反应迅速,但在调整其行为策略以适应新情况方面缺乏灵活性。基于这些假设,利用强化学习的计算理论,我们提出了两个过程之间仲裁的规范模型,该模型在搜索时间和决策准确性之间取得了近似最佳的平衡。在行为学上,该模型可以解释在学习的早期阶段行为对结果敏感的实验证据,但在后期阶段不敏感。它还解释说,当两个激励价值相等的选择同时可用时,行为仍然对结果敏感,即使经过广泛的培训。此外,该模型还可以解释学习过程中选择反应时的变化,以及随着选择次数的增加,反应时也增加的实验观察。在神经生物学上,通过假设中脑多巴胺神经元的时相和紧张性活动分别携带模型使用的奖赏预测误差和平均奖赏信号,该模型预测,虽然相性多巴胺通过加强刺激-反应关联间接影响行为,但紧张性多巴胺可以通过操纵习惯性系统和目标导向系统之间的竞争直接影响行为,从而影响反应时间。当面对不同的选择时,动物可以基于它们预先确定的习性做出反应,也可以通过考虑每个选择的短期和长期后果来做出反应。尽管习惯性的决策是快速的,但目标导向思维是一项耗时的任务。相反,习惯在被巩固后是不灵活的,但目标导向的决策可以在环境条件改变后迅速适应动物的策略。基于这两个决策系统的这些特点,我们提出了一种基于强化学习框架的计算模型,该模型在决策速度和行为灵活性之间取得了平衡。该模型的行为与观察结果一致,即在学习的早期阶段,动物的行为是以目标为导向的(灵活但缓慢),但在广泛学习后,它们的反应变得习惯性(不灵活,但速度快)。此外,该模型解释说,随着习惯性系统控制行为,动物的反应时间必须在学习过程中减少。该模型还将功能作用归因于多巴胺神经元在平衡习惯性系统和目标导向系统之间的竞争方面的紧张性活动。
Instrumental responses are hypothesized to be of two kinds: habitual and goal-directed, mediated by the sensorimotor and the associative cortico-basal ganglia circuits, respectively. The existence of the two heterogeneous associative learning mechanisms can be hypothesized to arise from the comparative advantages that they have at different stages of learning. In this paper, we assume that the goal-directed system is behaviourally flexible, but slow in choice selection. The habitual system, in contrast, is fast in responding, but inflexible in adapting its behavioural strategy to new conditions. Based on these assumptions and using the computational theory of reinforcement learning, we propose a normative model for arbitration between the two processes that makes an approximately optimal balance between search-time and accuracy in decision making. Behaviourally, the model can explain experimental evidence on behavioural sensitivity to outcome at the early stages of learning, but insensitivity at the later stages. It also explains that when two choices with equal incentive values are available concurrently, the behaviour remains outcome-sensitive, even after extensive training. Moreover, the model can explain choice reaction time variations during the course of learning, as well as the experimental observation that as the number of choices increases, the reaction time also increases. Neurobiologically, by assuming that phasic and tonic activities of midbrain dopamine neurons carry the reward prediction error and the average reward signals used by the model, respectively, the model predicts that whereas phasic dopamine indirectly affects behaviour through reinforcing stimulus-response associations, tonic dopamine can directly affect behaviour through manipulating the competition between the habitual and the goal-directed systems and thus, affect reaction time. When confronted with different alternatives, animals can respond either based on their pre-established habits, or by considering the short- and long-term consequences of each option. Whereas habitual decision making is fast, goal-directed thinking is a time-consuming task. Instead, habits are inflexible after being consolidated, but goal-directed decision making can rapidly adapt the animal's strategy after a change in environmental conditions. Based on these features of the two decision making systems, we suggest a computational model using the reinforcement learning framework, that makes a balance between the speed of decision making and behavioural flexibility. The behaviour of the model is consistent with the observation that at the early stages of learning, animals behave in a goal-directed way (flexible, but slow), but after extensive learning, their responses become habitual (inflexible, but fast). Moreover, the model explains that the animal's reaction time must decrease through the course of learning, as the habitual system takes control over behaviour. The model also attributes a functional role to the tonic activity of dopamine neurons in balancing the competition between the habitual and the goal-directed systems.
DOI: 10.3758/bf03342816
发表时间: 1964-01-01
期刊: PSYCHONOMIC SCIENCE
影响因子: --
作者:
ALLUISI, EA;STRAIN, GS;THURMOND, JB
通讯作者: THURMOND, JB
DOI: 10.1007/s11892-002-0065-7
发表时间: 2002-04-01
影响因子: 4.2
作者:
Evans, Mark L;Sherwin, Robert S
通讯作者: Sherwin, Robert S
DOI: 10.1111/j.2044-8295.1965.tb00944.x
发表时间: 1965-01-01
影响因子: 4
作者:
BROADBENT, DE;GREGORY, M
通讯作者: GREGORY, M
DOI: 10.1037/0097-7403.21.3.203
发表时间: 1995-07-01
期刊: JOURNAL OF EXPERIMENTAL PSYCHOLOGY-ANIMAL BEHAVIOR PROCESSES
影响因子: --
作者:
BALLEINE, BW;GARNER, C;DICKINSON, A
通讯作者: DICKINSON, A
DOI: 10.1080/14640748208400878
发表时间: 1982-01-01
期刊: QUARTERLY JOURNAL OF EXPERIMENTAL PSYCHOLOGY SECTION B-COMPARATIVE AND PHYSIOLOGICAL PSYCHOLOGY
影响因子: --
作者:
ADAMS, CD
通讯作者: ADAMS, CD