Structure learning in human sequential decision-making.

Structure learning in human sequential decision-making.
复制标题

DOI:
10.1371/journal.pcbi.1001003
复制
发表时间:
2010-12-02
影响因子:
4.3
通讯作者:
Schrater P
Schrater P
中科院分区:
生物学2区
文献类型:
--
作者:
Acuña DE;Schrater P

文献摘要

参考文献

被引文献

相似文献

对人类顺序决策的研究经常发现,与完美了解环境中奖励和事件如何产生的模型的理想行为者相比,人类的表现不是最优的。我们认为,人类面临的学习问题更复杂,因为它还涉及学习环境中奖励产生的结构,而不是次优问题。我们使用贝叶斯强化学习阐述了顺序决策任务中的结构学习问题,并表明学习奖励生成模型定性地改变了最优学习代理的行为。为了测试人们是否表现出结构学习,我们进行了涉及单臂和双臂强盗奖励模型的混合实验,其中结构学习产生了许多在以前的研究中被认为是次优的定性行为。我们的研究结果表明,人类可以以近乎最优的方式进行结构学习。每个决策实验都有一个结构,说明如何获得奖励,通常在实验开始时向受试者解释。参与者经常表现得好像他们理解实验结构一样,即使是在像决定他们应该从两个有偏差的硬币中选择哪一个来最大化产生“正面”的次数这样简单的任务中。我们假设参与者的行为不是由自上而下的指令驱动的——相反,参与者必须通过经验来学习奖励是如何产生的。我们使用完全合理的最优贝叶斯强化学习方法形式化这一假设,该方法模拟了顺序决策中的最优结构学习。在人类结构学习的实验测试中,我们表明人类以接近最优的方式从经验中学习奖励结构。我们的研究结果表明,行为表明人类是容易出错的,次优的决策者可以从最佳的学习方法中产生。我们的发现为以前被认为是非理性的行为,包括勘探不足和过度,提供了一个令人信服的新理性假设家族。
Studies of sequential decision-making in humans frequently find suboptimal performance relative to an ideal actor that has perfect knowledge of the model of how rewards and events are generated in the environment. Rather than being suboptimal, we argue that the learning problem humans face is more complex, in that it also involves learning the structure of reward generation in the environment. We formulate the problem of structure learning in sequential decision tasks using Bayesian reinforcement learning, and show that learning the generative model for rewards qualitatively changes the behavior of an optimal learning agent. To test whether people exhibit structure learning, we performed experiments involving a mixture of one-armed and two-armed bandit reward models, where structure learning produces many of the qualitative behaviors deemed suboptimal in previous studies. Our results demonstrate humans can perform structure learning in a near-optimal manner. Every decision-making experiment has a structure that specifies how rewards are obtained, which is usually explained to the subject at the beginning of the experiment. Participants frequently fail to act as if they understand the experimental structure, even in tasks as simple as determining which of two biased coins they should choose to maximize the number of trials that produce “heads”. We hypothesize that participants' behavior is not driven by top-down instructions—rather, participants must learn through experience how the rewards are generated. We formalize this hypothesis using a fully rational optimal Bayesian reinforcement learning approach that models optimal structure learning in sequential decision making. In an experimental test of structure learning in humans, we show that humans learn reward structure from experience in a near optimal manner. Our results demonstrate that behavior purported to show that humans are error-prone and suboptimal decision makers can result from an optimal learning approach. Our findings provide a compelling new family of rational hypotheses for behavior previously deemed irrational, including under- and over-exploration.
DOI: 10.1038/nature04766
发表时间: 2006-06-15
期刊: NATURE
影响因子: 64.8
作者:
Daw, Nathaniel D.;O'Doherty, John P.;Dayan, Peter;Seymour, Ben;Dolan, Raymond J.
通讯作者: Dolan, Raymond J.
DOI: 10.1037/h0046489
发表时间: 1962-01-01
影响因子: --
作者:
BRACKBILL, Y;BRAVOS, A
通讯作者: BRAVOS, A
在行动中构建学习。
DOI: 10.1016/j.bbr.2009.08.031
发表时间: 2010-01-20
影响因子: 2.7
作者:
Braun, Daniel A.;Mehring, Carsten;Wolpert, Daniel M.
通讯作者: Wolpert, Daniel M.
DOI: 10.1037/h0047727
发表时间: 1956-01-01
影响因子: --
作者:
EDWARDS, W
通讯作者: EDWARDS, W
DOI: 10.1016/j.conb.2010.02.008
发表时间: 2010-04
影响因子: 5.7
作者:
Gershman SJ;Niv Y
通讯作者: Niv Y