A reinforcement learning model of precommitment in decision making

A reinforcement learning model of precommitment in decision making
复制标题

DOI:
10.3389/fnbeh.2010.00184
复制
发表时间:
2010-01-01
影响因子:
3
通讯作者:
Redish, A. David
Redish, A. David
中科院分区:
医学3区
文献类型:
--
作者:
Kurth-Nelson, Zeb;Redish, A. David

文献摘要

被引文献

相似文献

成瘾和许多其他疾病都与冲动有关,当它立即可用时,次优选择是首选。冲动的一个解决方案是预先承诺:约束自己的未来,以避免被提供一个次优的选择。冲动的一种形式可以通过实验来测量,方法是在较早提供的较小奖励和较晚提供的较大奖励之间进行选择。冲动的被试更倾向于选择小而快的选择;然而,当提供一个预先承诺的选项时,即使是冲动的被试也会预先承诺大而晚的选择。是否预提交是在两个条件之间做出的决定:(A)最初的选择(更小-更快还是更大-更晚),以及(B)只有更大-更晚可用的新条件。人们已经观察到,预承诺出现的结果,固有的非指数延迟贴现的偏好逆转。在这里,我们表明,大多数模型的双曲贴现不能预承诺,但双曲贴现的分布式模型预承诺。使用这个模型,我们发现:(1)快折扣者可能比慢折扣者更有可能或更不可能预先承诺,这取决于预先承诺延迟;(2)对于一个恒定的小-早与大-晚偏好,较大奖励与较小奖励的较高比例增加了预先承诺的概率;(3)预先承诺对折扣曲线的形状高度敏感。这些预测意味着,操纵,改变折扣曲线,如饮食或上下文,可能会定性地影响预承诺。
Addiction and many other disorders are linked to impulsivity, where a suboptimal choice is preferred when it is immediately available. One solution to impulsivity is precommitment: constraining one's future to avoid being offered a suboptimal choice. A form of impulsivity can be measured experimentally by offering a choice between a smaller reward delivered sooner and a larger reward delivered later. Impulsive subjects are more likely to select the smaller-sooner choice; however, when offered an option to precommit, even impulsive subjects can precommit to the larger-later choice. To precommit or not is a decision between two conditions: (A) the original choice (smaller-sooner vs. larger-later), and (B) a new condition with only larger-later available. It has been observed that precommitment appears as a consequence of the preference reversal inherent in non-exponential delay-discounting. Here we show that most models of hyperbolic discounting cannot precommit, but a distributed model of hyperbolic discounting does precommit. Using this model, we find (1) faster discounters may be more or less likely than slow discounters to precommit, depending on the precommitment delay, (2) for a constant smaller-sooner vs. larger-later preference, a higher ratio of larger reward to smaller reward increases the probability of precommitment, and (3) precommitment is highly sensitive to the shape of the discount curve. These predictions imply that manipulations that alter the discount curve, such as diet or context, may qualitatively affect precommitment.