Parallel Representation of Value-Based and Finite State-Based Strategies in the Ventral and Dorsal Striatum.

Parallel Representation of Value-Based and Finite State-Based Strategies in the Ventral and Dorsal Striatum.
复制标题

DOI:
10.1371/journal.pcbi.1004540
复制
发表时间:
2015-11
影响因子:
4.3
通讯作者:
Doya K
Doya K
中科院分区:
生物学2区
文献类型:
--
作者:
Ito M;Doya K

文献摘要

被引文献

相似文献

以前的动物和人类行为学习的理论研究集中在二分法的价值为基础的策略,使用行动价值函数来预测奖励和基于模型的策略,使用内部模型来预测环境状态。然而,动物和人类经常采取简单的程序性行为,如“赢-留,输-转”策略,而没有明确的预测奖励或状态。在这里,我们考虑另一种策略,基于有限状态的策略,其中受试者根据其离散的内部状态选择动作,并根据所选择的动作和奖励结果更新状态。通过分析大鼠在自由选择任务中的选择行为,我们发现,基于有限状态的策略比基于价值和基于模型的策略更准确地拟合其行为选择。当拟合模型在同一任务下自主运行时,只有基于有限状态的策略才能再现选择序列的关键特征。记录从背外侧纹状体(DLS),背内侧纹状体(DMS),腹侧纹状体(VS)的神经活动的分析确定了显着分数的神经元在所有三个子区域的活动与个人状态的有限状态为基础的战略。在选择时的内部状态的信号被发现在DMS中,并为集群的状态被发现在VS中。此外,基于值的策略的动作值和状态值被编码在DMS和VS中,分别。这些结果表明,无论是基于价值的策略和基于有限状态的策略是在纹状体执行。决策的神经机制是在多种可能性中选择一种行动的认知过程,是神经科学的一个基本问题。先前的研究已经揭示了大脑皮层和基底神经节在决策中的作用,假设受试者采取基于价值的强化学习策略,其中每个候选行动的预期奖励被更新。然而,动物和人类经常使用简单的程序策略,如“赢-留,输-开关”。在这项研究中,我们考虑了一个有限的基于状态的策略,其中一个主体的行为取决于其离散的内部状态和更新的奖励反馈的基础上的状态。我们发现,在二元选择任务中,基于有限状态的策略比基于价值的策略能更好地再现大鼠的选择行为。有趣的是,纹状体中的神经元活动,一个基于奖励的学习的关键大脑区域,编码了关于这两种策略的信息。这些结果表明,无论是基于价值的策略和基于有限状态的策略是在纹状体执行。
Previous theoretical studies of animal and human behavioral learning have focused on the dichotomy of the value-based strategy using action value functions to predict rewards and the model-based strategy using internal models to predict environmental states. However, animals and humans often take simple procedural behaviors, such as the “win-stay, lose-switch” strategy without explicit prediction of rewards or states. Here we consider another strategy, the finite state-based strategy, in which a subject selects an action depending on its discrete internal state and updates the state depending on the action chosen and the reward outcome. By analyzing choice behavior of rats in a free-choice task, we found that the finite state-based strategy fitted their behavioral choices more accurately than value-based and model-based strategies did. When fitted models were run autonomously with the same task, only the finite state-based strategy could reproduce the key feature of choice sequences. Analyses of neural activity recorded from the dorsolateral striatum (DLS), the dorsomedial striatum (DMS), and the ventral striatum (VS) identified significant fractions of neurons in all three subareas for which activities were correlated with individual states of the finite state-based strategy. The signal of internal states at the time of choice was found in DMS, and for clusters of states was found in VS. In addition, action values and state values of the value-based strategy were encoded in DMS and VS, respectively. These results suggest that both the value-based strategy and the finite state-based strategy are implemented in the striatum. The neural mechanism of decision-making, a cognitive process to select one action among multiple possibilities, is a fundamental issue in neuroscience. Previous studies have revealed the roles of the cerebral cortex and the basal ganglia in decision-making, by assuming that subjects take a value-based reinforcement learning strategy, in which the expected reward for each action candidate is updated. However, animals and humans often use simple procedural strategies, such as “win-stay, lose-switch.” In this study, we consider a finite state-based strategy, in which a subject acts depending on its discrete internal state and updates the state based on reward feedback. We found that the finite state-based strategy could reproduce the choice behavior of rats in a binary choice task with higher accuracy than the value-based strategy. Interestingly, neuronal activity in the striatum, a crucial brain region for reward-based learning, encoded information regarding both strategies. These results suggest that both the value-based strategy and the finite state-based strategy are implemented in the striatum.