Momentary subjective well-being depends on learning and not reward.

Momentary subjective well-being depends on learning and not reward.
复制标题

短暂的主观幸福感取决于学习,而不是奖励。

DOI:
10.7554/elife.57977
复制
发表时间:
2020-11-17
期刊:
影响因子:
7.7
通讯作者:
Rutledge RB
Rutledge RB
中科院分区:
生物学1区
文献类型:
--
作者:
Blain B;Rutledge RB

文献摘要

被引文献

相似文献

主观的幸福或快乐往往与财富有关。最近的研究表明,短暂的快乐与奖励预测误差有关,这是经验和预测奖励之间的差异,是适应行为的关键组成部分。我们在强化学习任务中测试了受试者,其中奖励大小和概率不相关,这使我们能够将奖励和学习对幸福的贡献分开。使用计算模型,我们发现了稳定和不稳定的学习任务的收敛证据,即幸福感,就像行为一样,对学习相关变量(即概率预测误差)敏感。与行为不同,幸福对学习无关的变量(即奖励预测误差)不敏感。不断增加的波动性减少了过去的试验影响行为的数量,但不是幸福。最后,抑郁症状在不稳定的环境中比稳定的环境中更能降低幸福感。我们的研究结果表明,我们如何了解我们的世界可能比我们实际获得的奖励更重要。许多人相信,如果他们有更多的钱,他们会更快乐。而像中了彩票或获得大幅加薪这样的事件确实会让人们感到快乐,至少是暂时的。但最近的研究表明,在这种情况下,推动幸福感的主要因素并不是获得奖励的大小。相反,它是奖励与期望相匹配的程度。当你期望加薪1%时,得到10%的加薪会让你感到更快乐,而不是当你期望加薪20%时得到10%。预期奖励和实际奖励之间的这种差异被称为奖励预测误差。奖励预测错误在学习中起着关键作用。它们激励人们重复那些导致意外丰厚回报的行为。但它们也使人们能够更新他们对世界的信念,这本身就是有益的。奖励预测错误与幸福感有关,主要是因为它们帮助我们比以前更好地理解世界吗?为了验证这个想法,布兰和拉特利奇设计了一个任务,在这个任务中,获得奖励的可能性与奖励的大小无关。这项研究设计使得我们有可能将学习与奖励对每时每刻的幸福感的贡献分开。在这项任务中,志愿者必须决定两辆汽车中哪一辆会赢得比赛。在“稳定”状态下,其中一辆汽车总是有80%的获胜机会。在“不稳定”的条件下,一辆车有80%的机会赢得前20次试验。另一辆车则有80%的机会赢得接下来的20次试验。志愿者事先没有被告知这些概率,而是通过玩游戏来计算。然而,在每一次试验中,志愿者们都会看到如果他们选择了其中一辆汽车,并且那辆车继续获胜,他们将获得奖励。奖励的大小是随机变化的,与汽车获胜的可能性无关。每隔几次试验,志愿者们都会被要求在一个量表上指出他们目前的幸福程度。结果显示,志愿者在获胜后比失败后更快乐。平均而言,他们在稳定状态下比在波动状态下更快乐。这对于那些有抑郁症状的志愿者来说尤其如此。而且,胜利后的快乐并不取决于奖励有多大,而是取决于胜利后的惊喜程度。这些结果表明,对于我们的感受来说,我们如何了解周围的世界可能比我们直接获得的奖励更重要。在各种环境中测量幸福感可以帮助我们了解影响心理健康的因素。例如,目前的结果表明,不确定的环境可能对抑郁症患者特别不愉快。需要进一步的研究来理解为什么会出现这种情况。在真实的世界中,奖励往往是不确定的,也是不经常的,但学习可能仍然有潜力提高幸福感。
Subjective well-being or happiness is often associated with wealth. Recent studies suggest that momentary happiness is associated with reward prediction error, the difference between experienced and predicted reward, a key component of adaptive behaviour. We tested subjects in a reinforcement learning task in which reward size and probability were uncorrelated, allowing us to dissociate between the contributions of reward and learning to happiness. Using computational modelling, we found convergent evidence across stable and volatile learning tasks that happiness, like behaviour, is sensitive to learning-relevant variables (i.e. probability prediction error). Unlike behaviour, happiness is not sensitive to learning-irrelevant variables (i.e. reward prediction error). Increasing volatility reduces how many past trials influence behaviour but not happiness. Finally, depressive symptoms reduce happiness more in volatile than stable environments. Our results suggest that how we learn about our world may be more important for how we feel than the rewards we actually receive. Many people believe they would be happier if only they had more money. And events such as winning the lottery or receiving a large pay rise do make people happy, at least temporarily. But recent studies suggest that the main factor driving happiness on such occasions is not the size of the reward received. Instead, it is how well that reward matches up with expectations. Receiving a 10% pay rise when you were expecting 1% will make you feel happier than receiving 10% when you had been expecting 20%. This difference between an expected and an actual reward is referred to as a reward prediction error. Reward prediction errors have a key role in learning. They motivate people to repeat behaviours that led to unexpectedly large rewards. But they also enable people to update their beliefs about the world, which is rewarding in itself. Could it be that reward prediction errors are associated with happiness mainly because they help us understand the world a little better than before? To test this idea, Blain and Rutledge designed a task in which the likelihood of receiving a reward was unrelated to the size of the reward. This study design makes it possible to separate out the contributions of learning versus reward to moment-by-moment happiness. In the task, volunteers had to decide which of two cars would win a race. In the ‘stable’ condition, one of the cars always had an 80% chance of winning. In the ‘volatile’ condition, one car had an 80% chance of winning for the first 20 trials. The other car then had an 80% chance of winning for the next 20 trials. The volunteers were not told these probabilities in advance, but had to work them out by playing the game. However, on every trial, the volunteers were shown the reward they would receive if they chose either of the cars and that car went on to win. The size of the rewards varied at random and was unrelated to the likelihood of a car winning. Every few trials, the volunteers were asked to indicate their current level of happiness on a scale. The results showed that volunteers were happier after winning than after losing. On average they were also happier in the stable condition than in the volatile condition. This was especially true for volunteers with pre-existing symptoms of depression. Moreover, happiness after wins did not depend on how large the reward they got was, but instead simply on how surprised they were to win. These results suggest that how we learn about the world around us can be more important for how we feel than rewards we receive directly. Measuring happiness in various types of environment could help us understand factors affecting mental health. The current results suggest, for example, that uncertain environments may be especially unpleasant for people with depression. Further research is needed to understand why this might be the case. In the real world, rewards are often uncertain and infrequent, but learning may nevertheless have the potential to boost happiness.