Integrating reward information for prospective behaviour

Integrating reward information for prospective behaviour
复制标题

整合预期行为的奖励信息

DOI:
10.1101/2021.03.30.437719
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Hall-McMaster S
Hall-McMaster S
中科院分区:
--
文献类型:
--
作者:
Hall-McMaster S

文献摘要

相似文献

基于价值的决策通常是在静态的背景下研究的,参与者决定从现有的选项中选择哪个选项。然而,日常生活往往涉及一个额外的维度:决定何时选择最大化奖励。最近的证据表明,代理跟踪潜在的奖励的选项,更新其潜在的奖励估计的变化,以实现适当的选择时机(潜在的奖励跟踪)。然而,这种策略可能很难与提前估计最佳选择时间的策略区分开来,允许代理在选择之前等待预定的时间,而不需要监视选项的潜在奖励(距离目标跟踪)。在这里,我们表明,这些策略原则上可以分离。人类的大脑活动是用脑电图(EEG)记录的,而女性和男性参与者执行了一项新的决策任务。参与者被展示了一个选项,并决定何时选择它,因为它的潜在奖励随着试验的变化而变化。当潜在报酬未被提示时,它可以通过期权的初始价值和价值增长率的提示信息来估计。然后,我们使用代表性相似性分析(RSA)来评估EEG信号是否更接近潜在奖励跟踪或距离目标跟踪。这种方法成功地分离了任务中的策略。初始价值和增长率被转化为距离目标的信号,远远早于选择选项。潜在奖励无法独立解码。这些结果证明了使用高时间分辨率的神经记录来识别人脑中内部计算的决策变量的可行性。然而,外部世界并不总是告诉我们什么时候一个行动是最有益的,需要内部表征来指导行动时机。这种内部神经表征具有挑战性,因为它可能源于各种策略,其中许多策略对大脑活动做出类似的预测。本研究采用了一种新的方法来测试是否替代策略可以在原则上分离。使用代表性相似性分析(RSA),我们能够区分候选人的内部表示选择时间。这表明模式分析方法如何用于测量非侵入性神经数据中的潜在决策信息。
Value-based decision-making is often studied in a static context, where participants decide which option to select from those currently available. However, everyday life often involves an additional dimension: deciding when to select to maximize reward. Recent evidence suggests that agents track the latent reward of an option, updating changes in their latent reward estimate, to achieve appropriate selection timing (latent reward tracking). However, this strategy can be difficult to distinguish from one in which the optimal selection time is estimated in advance, allowing an agent to wait a predetermined amount of time before selecting, without needing to monitor an option's latent reward (distance-to-goal tracking). Here, we show that these strategies can in principle be dissociated. Human brain activity was recorded using electroencephalography (EEG), while female and male participants performed a novel decision task. Participants were shown an option and decided when to select it, as its latent reward changed from trial-to-trial. While the latent reward was uncued, it could be estimated using cued information about the option's starting value and value growth rate. We then used representational similarity analysis (RSA) to assess whether EEG signals more closely resembled latent reward tracking or distance-to-goal tracking. This approach successfully dissociated the strategies in this task. Starting value and growth rate were translated into a distance-to-goal signal, far in advance of selecting the option. Latent reward could not be independently decoded. These results demonstrate the feasibility of using high temporal resolution neural recordings to identify internally computed decision variables in the human brain.SIGNIFICANCE STATEMENTReward-seeking behavior involves acting at the right time. However, the external world does not always tell us when an action is most rewarding, necessitating internal representations that guide action timing. Specifying this internal neural representation is challenging because it might stem from a variety of strategies, many of which make similar predictions about brain activity. This study used a novel approach to test whether alternative strategies could be dissociated in principle. Using representational similarity analysis (RSA), we were able to distinguish between candidate internal representations for selection timing. This shows how pattern analysis methods can be used to measure latent decision information in noninvasive neural data.