Navigating complex decision spaces: Problems and paradigms in sequential choice.

Navigating complex decision spaces: Problems and paradigms in sequential choice.
复制标题

导航复杂的决策空间:顺序选择中的问题和范例。

DOI:
10.1037/a0033455
复制
发表时间:
2014-03
影响因子:
22.4
通讯作者:
Anderson, John R.
Anderson, John R.
中科院分区:
心理学1区
文献类型:
--
作者:
Walsh, Matthew M.;Anderson, John R.

文献摘要

参考文献

被引文献

相似文献

为了做出适应性行为,我们必须从我们行为的后果中学习。当行动的后果延迟发生时,这样做很困难。这就引入了时间信用分配的问题。当反馈遵循一系列决策时,个人应如何将功劳分配给构成该序列的中间行动?强化学习的研究为这个问题提供了两种通用的解决方案:无模型强化学习和基于模型的强化学习。在这篇综述中,我们研究了刺激响应和认知学习理论、习惯和目标导向控制以及无模型和基于模型的强化学习之间的联系。然后我们考虑一系列与时间学分分配相关的问题。这些包括二阶条件作用和二级强化物、潜在学习和绕行行为、部分可观察的马尔可夫决策过程、具有分布式结果的行动以及分层学习。我们询问人类和动物在面临这些问题时的行为方式是否与强化学习技术一致。自始至终,我们都致力于确定无模型和基于模型的强化学习的神经基础。前一类技术是根据神经递质多巴胺及其对基底神经节的影响来理解的。后者可以理解为包括前额皮质、小脑内侧颞叶和基底神经节的分布式网络。强化学习技术不仅对人类和动物行为有自然的解释,而且还为理解神经奖励评估和​​行动选择提供了有用的框架。
To behave adaptively, we must learn from the consequences of our actions. Doing so is difficult when the consequences of an action follow a delay. This introduces the problem of temporal credit assignment. When feedback follows a sequence of decisions, how should the individual assign credit to the intermediate actions that comprise the sequence? Research in reinforcement learning provides two general solutions to this problem: model-free reinforcement learning and model-based reinforcement learning. In this review, we examine connections between stimulus-response and cognitive learning theories, habitual and goal-directed control, and model-free and model-based reinforcement learning. We then consider a range of problems related to temporal credit assignment. These include second-order conditioning and secondary reinforcers, latent learning and detour behavior, partially observable Markov decision processes, actions with distributed outcomes, and hierarchical learning. We ask whether humans and animals, when faced with these problems, behave in a manner consistent with reinforcement learning techniques. Throughout, we seek to identify neural substrates of model-free and model-based reinforcement learning. The former class of techniques is understood in terms of the neurotransmitter dopamine and its effects in the basal ganglia. The latter is understood in terms of a distributed network of regions including the prefrontal cortex, medial temporal lobes cerebellum, and basal ganglia. Not only do reinforcement learning techniques have a natural interpretation in terms of human and animal behavior, but they also provide a useful framework for understanding neural reward valuation and action selection.
DOI: 10.1016/j.neuroimage.2006.01.001
发表时间: 2006-06-01
期刊: NEUROIMAGE
影响因子: 5.7
作者:
Abler, Birgit;Walter, Henrik;Spitzer, Manfred
通讯作者: Spitzer, Manfred
DOI: 10.1080/14640748108400816
发表时间: 1981-01-01
期刊: QUARTERLY JOURNAL OF EXPERIMENTAL PSYCHOLOGY SECTION B-COMPARATIVE AND PHYSIOLOGICAL PSYCHOLOGY
影响因子: --
作者:
ADAMS, CD;DICKINSON, A
通讯作者: DICKINSON, A
DOI: 10.1037/0097-7403.21.3.203
发表时间: 1995-07-01
期刊: JOURNAL OF EXPERIMENTAL PSYCHOLOGY-ANIMAL BEHAVIOR PROCESSES
影响因子: --
作者:
BALLEINE, BW;GARNER, C;DICKINSON, A
通讯作者: DICKINSON, A
DOI: 10.1016/j.cognition.2008.08.011
发表时间: 2009-12
期刊: Cognition
影响因子: 3.4
作者:
Botvinick MM;Niv Y;Barto AG
通讯作者: Barto AG
DOI: 10.1152/jn.01030.2009
发表时间: 2010-05-01
影响因子: 2.5
作者:
Bray, Signe;Shimojo, Shinsuke;O'Doherty, John P.
通讯作者: O'Doherty, John P.