Simple Plans or Sophisticated Habits? State, Transition and Learning Interactions in the Two-Step Task.

Simple Plans or Sophisticated Habits? State, Transition and Learning Interactions in the Two-Step Task.
复制标题

DOI:
10.1371/journal.pcbi.1004648
复制
发表时间:
2015-12
影响因子:
4.3
通讯作者:
Dayan P
Dayan P
中科院分区:
生物学2区
文献类型:
--
作者:
Akam T;Costa R;Dayan P

文献摘要

被引文献

相似文献

最近开发的“两步”行为任务有望区分基于模型的强化学习和无模型强化学习,同时生成具有决策变量参数变化的神经生理学友好的决策数据集。这些令人满意的特点促使其得到广泛采用。在这里,我们分析了一系列不同的策略和结构的过渡和结果之间的相互作用,以检查可以从行为表现中学到什么的限制。这项任务涉及到对随机性的需要(允许策略被区分)和对确定性的需要之间的权衡,因此值得受试者投入努力以最佳地利用突发事件。我们通过模拟表明,在一定条件下,无模型策略可以伪装成基于模型的。我们首先表明,看似无害的修改任务结构可以诱导在审判开始时的动作值和随后的审判事件之间的相关性,这样的方式,分析的基础上比较连续的试验可能会导致错误的结论。我们确认的权力,建议的修正分析,可以缓解这个问题。然后,我们考虑无模型强化学习策略,该策略利用获得奖励的位置与具有高期望值的动作之间的相关性。在这些更复杂的分析下,这些行为似乎是基于模型的。开发两步任务作为行为神经科学工具的全部潜力需要理解这些问题。规划是使用行动后果的预测模型来指导决策。规划在人类行为中起着至关重要的作用,但隔离它的贡献是具有挑战性的,因为它是由控制系统补充的,控制系统直接从强化的历史中学习行为的价值,导致从状态到行为的自动映射,通常称为习惯。我们的研究探讨了最近开发的行为任务,使用多步决策树的选择,以区分规划基于价值的控制。我们使用模拟比较了各种策略,显示了产生类似于计划的行为的范围,但实际上是从特定类型的状态到行动的固定映射。这些结果表明,当一个规划问题是反复面临的,复杂的自动化策略可能会被开发出来,它确定,事实上有一个有限数量的相关国家的世界每个适当的固定或习惯的反应。理解这些策略对于设计和解释旨在隔离计划对行为的贡献的任务是重要的。这些策略也具有独立的科学意义,因为它们可能有助于在复杂环境中实现行为自动化。
The recently developed ‘two-step’ behavioural task promises to differentiate model-based from model-free reinforcement learning, while generating neurophysiologically-friendly decision datasets with parametric variation of decision variables. These desirable features have prompted its widespread adoption. Here, we analyse the interactions between a range of different strategies and the structure of transitions and outcomes in order to examine constraints on what can be learned from behavioural performance. The task involves a trade-off between the need for stochasticity, to allow strategies to be discriminated, and a need for determinism, so that it is worth subjects’ investment of effort to exploit the contingencies optimally. We show through simulation that under certain conditions model-free strategies can masquerade as being model-based. We first show that seemingly innocuous modifications to the task structure can induce correlations between action values at the start of the trial and the subsequent trial events in such a way that analysis based on comparing successive trials can lead to erroneous conclusions. We confirm the power of a suggested correction to the analysis that can alleviate this problem. We then consider model-free reinforcement learning strategies that exploit correlations between where rewards are obtained and which actions have high expected value. These generate behaviour that appears model-based under these, and also more sophisticated, analyses. Exploiting the full potential of the two-step task as a tool for behavioural neuroscience requires an understanding of these issues. Planning is the use of a predictive model of the consequences of actions to guide decision making. Planning plays a critical role in human behaviour, but isolating its contribution is challenging because it is complemented by control systems which learn values of actions directly from the history of reinforcement, resulting in automatized mappings from states to actions often termed habits. Our study examined a recently developed behavioural task which uses choices in a multi-step decision tree to differentiate planning from value-based control. We compared various strategies using simulations, showing a range that produce behaviour that resembles planning but in fact arises as a fixed mapping from particular sorts of states to action. These results show that when a planning problem is faced repeatedly, sophisticated automatization strategies may be developed which identify that there are in fact a limited number of relevant states of the world each with an appropriate fixed or habitual response. Understanding such strategies is important for the design and interpretation of tasks which aim to isolate the contribution of planning to behaviour. Such strategies are also of independent scientific interest as they may contribute to automatization of behaviour in complex environments.