Humans can adopt optimal discounting strategy under real-time constraints.

Humans can adopt optimal discounting strategy under real-time constraints.
复制标题

DOI:
10.1371/journal.pcbi.0020152
复制
发表时间:
2006-11-10
影响因子:
4.3
通讯作者:
Doya K
Doya K
中科院分区:
生物学2区
文献类型:
--
作者:
Schweighofer N;Shishida K;Han CE;Okamoto Y;Tanaka SC;Yamawaki S;Doya K

文献摘要

参考文献

被引文献

相似文献

我们每天在较大的延迟奖励和较小的更直接的奖励之间做出选择的关键是,随时间折扣奖励的函数的形状和陡峭度。尽管人工智能的研究倾向于在不确定的环境中采用指数折现,但对人类和动物的研究一直显示出双曲线折现。我们研究了人类在具有时间限制的奖励决策任务中的表现,其中每个选择都会影响后续试验的剩余时间,并且每次试验的延迟都是不同的。在这个实验中,我们证明了大多数被试都采用了指数贴现。此外,我们通过分析证实了指数折扣,其衰减率与我们的实验对象使用的指数折扣相当,在我们的任务中最大化了总奖励收益。我们的研究结果表明,时间折扣的特定形状和陡峭程度是由受试者面临的任务决定的,并质疑双曲奖励折扣作为普遍原则的概念。当我们在两个选项之间做出选择时,我们比较它们的结果值,并选择值较大的选项。然而,如果一种选择导致更大的延迟奖励,而另一种选择导致更小的即时奖励呢?当然,我们会为更大的奖励分配更大的价值,但如果奖励要晚一些才发放,那就会“打折扣”。因此,该值是延迟的单调递减函数。先前的行为研究一再证明,人类和动物对延迟奖励的低估是夸张的。这具有实际意义,因为双曲贴现有时会导致“非理性”的偏好逆转:例如,一个人可能更喜欢51天后的两个苹果,而不是50天后的一个苹果,但如果日期更近,他更喜欢今天的一个苹果,而不是明天的两个苹果。相反,指数贴现总是“合理的”,因为它预测了不变的偏好。在这里,在一个模仿动物觅食的新任务中,施威弗和他的同事们发现,人类也能以指数方式低估奖励。此外,值得注意的是,通过采用指数贴现,他们的受试者最大限度地提高了他们的总收益。因此,根据手头的任务,作者的研究表明,人类可以灵活地选择奖励折扣的类型,并且可以表现出最大化长期收益的理性行为。
Critical to our many daily choices between larger delayed rewards, and smaller more immediate rewards, are the shape and the steepness of the function that discounts rewards with time. Although research in artificial intelligence favors exponential discounting in uncertain environments, studies with humans and animals have consistently shown hyperbolic discounting. We investigated how humans perform in a reward decision task with temporal constraints, in which each choice affects the time remaining for later trials, and in which the delays vary at each trial. We demonstrated that most of our subjects adopted exponential discounting in this experiment. Further, we confirmed analytically that exponential discounting, with a decay rate comparable to that used by our subjects, maximized the total reward gain in our task. Our results suggest that the particular shape and steepness of temporal discounting is determined by the task that the subject is facing, and question the notion of hyperbolic reward discounting as a universal principle. When we make a choice between two options, we compare the values of their outcomes and select the option with a larger value. However, what if one option leads to a larger delayed reward and the other leads to a smaller more immediate reward? Naturally, we assign a larger value for a larger reward, but it is “discounted” if the reward is to be delivered later. Thus, the value is a monotonically decreasing function of the delays. Previous behavioral studies have repeatedly demonstrated that humans and animals discount delayed rewards hyperbolically. This has practical importance, as hyperbolic discounting can sometimes lead to “irrational” preference reversal: for instance, an individual may prefer two apples in 51 days to one apple in 50 days, but if the days come closer, he prefers one apple today to two apples tomorrow. On the contrary, exponential discounting is always “rational,” as it predicts constant preference. Here, in a new task that mimics animal foraging, and that uses delayed monetary rewards, Schweighofer and colleagues showed that humans can also discount reward exponentially. Furthermore, it is remarkable that by adopting exponential discounting, their subjects maximized their total gain. Thus, depending on the task at hand, the authors' study suggests that humans can flexibly choose the type of reward discounting, and can exhibit rational behavior that maximizes long-term gains.
DOI: 10.1901/jeab.1995.64-263
发表时间: 1995-11-01
影响因子: 2.7
作者:
MYERSON, J;GREEN, L
通讯作者: GREEN, L
DOI: 10.1017/s0140525x05000117
发表时间: 2005-10-01
影响因子: 29.3
作者:
Ainslie, G
通讯作者: Ainslie, G
DOI: 10.1037/1064-1297.5.3.256
发表时间: 1997-08-01
影响因子: 2.3
作者:
Madden, GJ;Petry, NM;Bickel, WK
通讯作者: Bickel, WK
DOI: 10.1006/obhd.1995.1086
发表时间: 1995-10-01
影响因子: 4.6
作者:
KIRBY, KN;MARAKOVIC, NN
通讯作者: MARAKOVIC, NN
DOI: 10.1093/beheco/7.3.341
发表时间: 1996-09-01
期刊: BEHAVIORAL ECOLOGY
影响因子: 2.4
作者:
Bateson, M;Kacelnik, A
通讯作者: Kacelnik, A