The nature of belief-directed exploratory choice in human decision-making.

The nature of belief-directed exploratory choice in human decision-making.
复制标题

DOI:
10.3389/fpsyg.2011.00398
复制
发表时间:
2011
影响因子:
3.8
通讯作者:
Love BC
Love BC
中科院分区:
心理学3区
文献类型:
--
作者:
Knox WB;Otto AR;Stone P;Love BC

文献摘要

被引文献

相似文献

在非平稳环境中,利用当前有利的选项和通过探索过去被证明回报较低的鲜为人知的选项来获取信息之间存在冲突。这些任务中的最佳决策需要考虑环境的未来状态(即,规划),并在观察与选择相关的结果后,适当地更新关于环境状态的信念。最佳信念更新是反射性的,因为信念可以在不直接观察环境变化的情况下改变。例如,在10秒过去之后,人们可能正确地认为最后观察到的交通灯是红色,现在更可能是绿色。为了理解人类决策时,奖励与选择选项随着时间的推移而变化,我们开发了一个变种的经典的“土匪”任务,既丰富到足以涵盖相关的现象,并足够容易处理,以允许理想的演员分析顺序选择行为。我们评估人们是否会自反性地更新对环境状态的信念(即,仅响应于观察到的奖励结构的变化)或反思的方式。与纯粹的“随机”探索行为的帐户,基于模型的分析主题的选择和lavelet表明,人们是反思的信念更新。然而,与理想行动者模型不同,我们的分析表明,人们的选择行为并不反映对未来环境状态的考虑。因此,尽管人们以与理想行动者一致的反思方式更新信念,但他们并没有进行最佳的长期规划,而是在每次试验中都短视地选择被认为具有最高即时回报的选项。
In non-stationary environments, there is a conflict between exploiting currently favored options and gaining information by exploring lesser-known options that in the past have proven less rewarding. Optimal decision-making in such tasks requires considering future states of the environment (i.e., planning) and properly updating beliefs about the state of the environment after observing outcomes associated with choices. Optimal belief-updating is reflective in that beliefs can change without directly observing environmental change. For example, after 10 s elapse, one might correctly believe that a traffic light last observed to be red is now more likely to be green. To understand human decision-making when rewards associated with choice options change over time, we develop a variant of the classic “bandit” task that is both rich enough to encompass relevant phenomena and sufficiently tractable to allow for ideal actor analysis of sequential choice behavior. We evaluate whether people update beliefs about the state of environment in a reflexive (i.e., only in response to observed changes in reward structure) or reflective manner. In contrast to purely “random” accounts of exploratory behavior, model-based analyses of the subjects’ choices and latencies indicate that people are reflective belief updaters. However, unlike the Ideal Actor model, our analyses indicate that people’s choice behavior does not reflect consideration of future environmental states. Thus, although people update beliefs in a reflective manner consistent with the Ideal Actor, they do not engage in optimal long-term planning, but instead myopically choose the option on every trial that is believed to have the highest immediate payoff.