Risk, unexpected uncertainty, and estimation uncertainty: Bayesian learning in unstable settings.

Risk, unexpected uncertainty, and estimation uncertainty: Bayesian learning in unstable settings.
复制标题

DOI:
10.1371/journal.pcbi.1001048
复制
发表时间:
2011-01-20
影响因子:
4.3
通讯作者:
Bossaerts P
Bossaerts P
中科院分区:
生物学2区
文献类型:
--
作者:
Payzan-LeNestour E;Bossaerts P

文献摘要

参考文献

被引文献

相似文献

最近,有证据表明,在一个六臂不安分的强盗问题中,人类使用贝叶斯更新而不是(无模型)强化算法进行学习。在这里,我们调查这对人类欣赏不确定性意味着什么。在我们的任务中,一名贝叶斯学习者区分了三种同样显著的不确定性水平。首先,贝叶斯感知到了不可减少的不确定性或风险:即使知道一只给定手臂的回报概率,结果也是不确定的。其次,存在(参数)估计的不确定性或模糊性:收益概率未知,需要估计。第三,武器变化的结果概率:突然的跳跃被称为意想不到的不确定性。我们记录了在实验过程中这三种水平的不确定性是如何演变的,以及它是如何影响学习速度的。然后,我们放大了估计的不确定性,这被认为是探索的动力,尽管有证据表明人们普遍厌恶模棱两可。我们的数据证实了后者。我们讨论了神经证据,这些证据预示了人类区分三种不确定程度的能力。最后,我们调查了实施贝叶斯学习的人类能力的边界。我们用不同的指令重复实验,反映了不同程度的结构不确定性。在不确定性的第四个概念下,贝叶斯更新并不比(无模型的)强化学习更好地解释选择。退出问卷显示,参与者仍然没有意识到存在意想不到的不确定性,也没有获得用于实施贝叶斯更新的正确模型。人类学习不断变化的或有回报的能力意味着,他们至少感知到三个级别的不确定性:风险,这反映了即使在一切都知道之后也不完全的远见;(参数)估计的不确定性,即关于结果概率的不确定性;以及意外的不确定性,即概率的突然变化。我们描述了这些不确定性水平是如何在自然采样任务中演变的,在该任务中,人类的选择可靠地反映了最优(贝叶斯)学习,以及它们的演变如何改变了学习率。然后,我们放大估计的不确定性。感知估计不确定性(也称为模糊性)的能力是一种优点,因为除了允许一个人以最佳方式学习之外,它可能会引导更有效的探索;但对估计不确定性的厌恶可能是不适应的。在这里,我们表明,参与者的选择反映了对估计不确定性的厌恶。我们讨论了过去的成像研究如何预示了人类区分不同不确定性概念的能力。此外,我们还证明,参与者进行这种区分的能力依赖于对回报生成模型的充分揭示。当我们引入结构不确定性时,参与者没有意识到我们任务中的跳跃,而是回到了无模型强化学习。
Recently, evidence has emerged that humans approach learning using Bayesian updating rather than (model-free) reinforcement algorithms in a six-arm restless bandit problem. Here, we investigate what this implies for human appreciation of uncertainty. In our task, a Bayesian learner distinguishes three equally salient levels of uncertainty. First, the Bayesian perceives irreducible uncertainty or risk: even knowing the payoff probabilities of a given arm, the outcome remains uncertain. Second, there is (parameter) estimation uncertainty or ambiguity: payoff probabilities are unknown and need to be estimated. Third, the outcome probabilities of the arms change: the sudden jumps are referred to as unexpected uncertainty. We document how the three levels of uncertainty evolved during the course of our experiment and how it affected the learning rate. We then zoom in on estimation uncertainty, which has been suggested to be a driving force in exploration, in spite of evidence of widespread aversion to ambiguity. Our data corroborate the latter. We discuss neural evidence that foreshadowed the ability of humans to distinguish between the three levels of uncertainty. Finally, we investigate the boundaries of human capacity to implement Bayesian learning. We repeat the experiment with different instructions, reflecting varying levels of structural uncertainty. Under this fourth notion of uncertainty, choices were no better explained by Bayesian updating than by (model-free) reinforcement learning. Exit questionnaires revealed that participants remained unaware of the presence of unexpected uncertainty and failed to acquire the right model with which to implement Bayesian updating. The ability of humans to learn changing reward contingencies implies that they perceive, at a minimum, three levels of uncertainty: risk, which reflects imperfect foresight even after everything is learned; (parameter) estimation uncertainty, i.e., uncertainty about outcome probabilities; and unexpected uncertainty, or sudden changes in the probabilities. We describe how these levels of uncertainty evolve in a natural sampling task in which human choices reliably reflect optimal (Bayesian) learning, and how their evolution changes the learning rate. We then zoom in on estimation uncertainty. The ability to sense estimation uncertainty (also known as ambiguity) is a virtue because, besides allowing one to learn optimally, it may guide more effective exploration; but aversion to estimation uncertainty may be maladaptive. Here, we show that participant choices reflected aversion to estimation uncertainty. We discuss how past imaging studies foreshadowed the ability of humans to distinguish the different notions of uncertainty. Also, we document that the ability of participants to do such distinction relies on sufficient revelation of the payoff-generating model. When we induced structural uncertainty, participants did not gain awareness of the jumps in our task, and fell back to model-free reinforcement learning.
DOI: 10.1038/nature04766
发表时间: 2006-06-15
期刊: NATURE
影响因子: 64.8
作者:
Daw, Nathaniel D.;O'Doherty, John P.;Dayan, Peter;Seymour, Ben;Dolan, Raymond J.
通讯作者: Dolan, Raymond J.
DOI: 10.1098/rstb.2007.2098
发表时间: 2007-05-29
影响因子: 6.3
作者:
Cohen, Jonathan D.;McClure, Samuel M.;Yu, Angela J.
通讯作者: Yu, Angela J.
DOI: 10.1111/j.2517-6161.1995.tb02015.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
DRAPER, D
通讯作者: DRAPER, D
DOI: 10.1016/j.tics.2006.05.004
发表时间: 2006-07-01
影响因子: 19.9
作者:
Courville, Aaron C.;Daw, Nathaniel D.;Touretzky, David S.
通讯作者: Touretzky, David S.
DOI: 10.1126/science.1077349
发表时间: 2003-03-21
期刊: SCIENCE
影响因子: 56.9
作者:
Fiorillo, CD;Tobler, PN;Schultz, W
通讯作者: Schultz, W