A Reinforcement Learning Approach to Understanding Procrastination: Does Inaccurate Value Approximation Cause Irrational Postponing of a Task?

A Reinforcement Learning Approach to Understanding Procrastination: Does Inaccurate Value Approximation Cause Irrational Postponing of a Task?
复制标题

DOI:
10.3389/fnins.2021.660595
复制
发表时间:
2021
影响因子:
4.3
通讯作者:
Morita K
Morita K
中科院分区:
医学2区
文献类型:
--
作者:
Feng Z;Nagase AM;Morita K

文献摘要

参考文献

相似文献

拖延是一种自愿但非理性地推迟一项任务的行为,尽管意识到拖延会导致更糟糕的后果。从影响因素到理论模型,心理学界对它进行了广泛的研究。从价值决策和强化学习的角度来看,拖延是由认知限制导致的非最优选择造成的。然而,究竟是什么样的认知局限性所涉及的,仍然是难以捉摸的。在目前的研究中,我们研究了一种特定类型的认知限制,即不充分的状态表征导致的不准确的评估,是否会导致拖延。最近的研究表明,人类可能会采用一种特殊类型的状态表示,称为后继表示(SR),人类可以学习用相对低维的特征来表示状态。结合这些建议,我们假设了SR的降维版本。我们模拟了一个“学生”在学期中做作业的一系列行为,当推迟做作业时(即,拖延)是不允许的,而在假期中,何时拖延或不拖延可以自由选择。我们假设,在没有拖延的策略下,“学生”已经获得了与完成作业的每一步相对应的每个状态的刚性缩减SR。“学生”通过时间差(TD)学习学习每个状态的近似值,该近似值被计算为刚性简化SR中状态特征的线性函数。在假期中,“学生”在每个时间步根据这些近似值决定是否拖延。仿真结果表明,减少SR为基础的RL模型产生拖延行为,恶化跨事件。根据“学生”的近似价值观,拖延是更好的选择,而根据真实价值观,不拖延大多更好。因此,目前的模型产生的拖延行为所造成的不准确的值近似,这是由于采用减少SR作为状态表示。这些发现表明,减少SR,或者更一般地说,状态表征的维度减少,可能是导致拖延的认知限制的一种潜在形式。
Procrastination is the voluntary but irrational postponing of a task despite being aware that the delay can lead to worse consequences. It has been extensively studied in psychological field, from contributing factors, to theoretical models. From value-based decision making and reinforcement learning (RL) perspective, procrastination has been suggested to be caused by non-optimal choice resulting from cognitive limitations. Exactly what sort of cognitive limitations are involved, however, remains elusive. In the current study, we examined if a particular type of cognitive limitation, namely, inaccurate valuation resulting from inadequate state representation, would cause procrastination. Recent work has suggested that humans may adopt a particular type of state representation called the successor representation (SR) and that humans can learn to represent states by relatively low-dimensional features. Combining these suggestions, we assumed a dimension-reduced version of SR. We modeled a series of behaviors of a “student” doing assignments during the school term, when putting off doing the assignments (i.e., procrastination) is not allowed, and during the vacation, when whether to procrastinate or not can be freely chosen. We assumed that the “student” had acquired a rigid reduced SR of each state, corresponding to each step in completing an assignment, under the policy without procrastination. The “student” learned the approximated value of each state which was computed as a linear function of features of the states in the rigid reduced SR, through temporal-difference (TD) learning. During the vacation, the “student” made decisions at each time-step whether to procrastinate based on these approximated values. Simulation results showed that the reduced SR-based RL model generated procrastination behavior, which worsened across episodes. According to the values approximated by the “student,” to procrastinate was the better choice, whereas not to procrastinate was mostly better according to the true values. Thus, the current model generated procrastination behavior caused by inaccurate value approximation, which resulted from the adoption of the reduced SR as state representation. These findings indicate that the reduced SR, or more generally, the dimension reduction in state representation, can be a potential form of cognitive limitation that leads to procrastination.
DOI: 10.1037/a0037015
发表时间: 2014-07-01
影响因子: 5.4
作者:
Collins, Anne G. E.;Frank, Michael J.
通讯作者: Frank, Michael J.
DOI: 10.1098/rspb.2018.1645
发表时间: 2018-11-21
影响因子: 4.7
作者:
Gardner, Matthew P. H.;Schoenbaum, Geoffrey;Gershman, Samuel J.
通讯作者: Gershman, Samuel J.
DOI: 10.3389/fncom.2013.00174
发表时间: 2013-12-06
影响因子: 3.2
作者:
Helie S;Chakravarthy S;Moustafa AA
通讯作者: Moustafa AA
中断多巴胺对未来奖励编码的可取消成本和利益编码。
DOI: 10.1038/nn.2460
发表时间: 2010-01
影响因子: 25
作者:
Gan JO;Walton ME;Phillips PE
通讯作者: Phillips PE
DOI: 10.1523/jneurosci.4515-08.2009
发表时间: 2009-04-08
期刊: The Journal of neuroscience : the official journal of the Society for Neuroscience
影响因子: --
作者:
Croxson PL;Walton ME;O'Reilly JX;Behrens TE;Rushworth MF
通讯作者: Rushworth MF