Habits, action sequences and reinforcement learning.

Habits, action sequences and reinforcement learning.
复制标题

DOI:
10.1111/j.1460-9568.2012.08050.x
复制
发表时间:
2012-04
期刊:
The European journal of neuroscience
影响因子:
--
通讯作者:
Balleine BW
Balleine BW
中科院分区:
其他
文献类型:
--
作者:
Dezfouli A;Balleine BW

文献摘要

被引文献

相似文献

现在人们普遍认为,工具性行为可以是目标导向的,也可以是习惯性的;前者是迅速获得的,并由其结果调节,后者是反射性的,由先行刺激而不是其后果引起。基于模型的强化学习(RL)为目标导向的行动提供了一个优雅的描述。通过接触状态、行动和奖励,智能体迅速构建了一个世界模型,并可以根据环境和评估需求的相当抽象的变化选择适当的行动。这个模型是强大的,但有一个问题,解释习惯性行为的发展。为了解释习惯,理论家们认为需要另一个动作控制器,称为无模型强化学习,它不形成世界的模型,而是将动作值缓存在状态中,允许状态基于其奖励历史而不是其后果来选择动作。然而,模型的重要预测仍然存在一些问题,最明显的是无模型强化学习无法正确预测习惯性行为对行为奖励偶然性变化的不敏感性。在这里,我们建议,引入无模型的RL在工具条件反射是不必要的,并表明,重新概念化的习惯作为动作序列,允许基于模型的RL被应用到目标导向和习惯性的行动的方式符合什么真实的动物做。这种方法对目前研究习惯的方式具有重要意义,并产生了新的实验预测。
It is now widely accepted that instrumental actions can be either goal-directed or habitual; whereas the former are rapidly acquire and regulated by their outcome, the latter are reflexive, elicited by antecedent stimuli rather than their consequences. Model-based reinforcement learning (RL) provides an elegant description of goal-directed action. Through exposure to states, actions and rewards, the agent rapidly constructs a model of the world and can choose an appropriate action based on quite abstract changes in environmental and evaluative demands. This model is powerful but has a problem explaining the development of habitual actions. To account for habits, theorists have argued that another action controller is required, called model-free RL, that does not form a model of the world but rather caches action values within states allowing a state to select an action based on its reward history rather than its consequences. Nevertheless, there are persistent problems with important predictions from the model; most notably the failure of model-free RL correctly to predict the insensitivity of habitual actions to changes in the action-reward contingency. Here, we suggest that introducing model-free RL in instrumental conditioning is unnecessary and demonstrate that reconceptualizing habits as action sequences allows model-based RL to be applied to both goal-directed and habitual actions in a manner consistent with what real animals do. This approach has significant implications for the way habits are currently investigated and generates new experimental predictions.