Optimism and pessimism in optimised replay.

Optimism and pessimism in optimised replay.
复制标题

DOI:
10.1371/journal.pcbi.1009634
复制
发表时间:
2022-01
影响因子:
4.3
通讯作者:
Dayan P
Dayan P
中科院分区:
生物学2区
文献类型:
--
作者:
Antonov G;Gagne C;Eldar E;Dayan P

文献摘要

参考文献

相似文献

众所周知,任务相关轨迹的重播有助于巩固记忆和提高任务绩效。大量的实验数据表明,重放序列的内容具有高度的特殊性,并且可以受到奖赏和其他显著任务变量的影响。然而,控制重播序列选择的规则仍然知之甚少。最近的一个理论建议是,在决策问题中,重演经验的优先顺序是基于它们对行动选择的影响。我们表明,这意味着,当受试者由于健忘而在他们的内部世界模型中经历了大量的不确定性时,他们应该重播他们功能失调选择的次优行为,而不是最优的行为。我们用这一点来解释最近的实验数据,这些数据准确地证明了悲观的重播,将模型参数与个别受试者的选择相匹配。当动物睡着或完全清醒时,它们大脑中的神经元群体概括了与延长的与行为相关的经验相关的活动。这个过程被称为重演,在啮齿类动物身上已经确立了很长一段时间,最近在人类身上,它对决策任务的良好表现很重要。在这些时代重演的具体经验遵循高度有序的模式,但确定其优先顺序的机制仍未完全理解。一个有希望的理论建议是,每一次重演经历的选择都是以这样一种方式进行的,即随后的学习对动物随后的表现最有帮助。最近的一项研究报告了一个令人惊讶的结果:在计划任务中取得高表现的人倾向于重演他们发现不是最优的行动,这与在随后的表现中有用地反对这些行动有关。在这项研究中,我们检查了这种悲观的重播形式的性质,并表明它完全适合健忘的代理人。我们分析了遗忘在模型重播选择中的作用,并使用人类受试者数据验证了我们的预测。
The replay of task-relevant trajectories is known to contribute to memory consolidation and improved task performance. A wide variety of experimental data show that the content of replayed sequences is highly specific and can be modulated by reward as well as other prominent task variables. However, the rules governing the choice of sequences to be replayed still remain poorly understood. One recent theoretical suggestion is that the prioritization of replay experiences in decision-making problems is based on their effect on the choice of action. We show that this implies that subjects should replay sub-optimal actions that they dysfunctionally choose rather than optimal ones, when, by being forgetful, they experience large amounts of uncertainty in their internal models of the world. We use this to account for recent experimental data demonstrating exactly pessimal replay, fitting model parameters to the individual subjects’ choices. When animals are asleep or restfully awake, populations of neurons in their brains recapitulate activity associated with extended behaviourally-relevant experiences. This process is called replay, and it has been established for a long time in rodents, and very recently in humans, to be important for good performance in decision-making tasks. The specific experiences which are replayed during those epochs follow highly ordered patterns, but the mechanisms which establish their priority are still not fully understood. One promising theoretical suggestion is that each replay experience is chosen in such a way that the learning that ensues is most helpful for the subsequent performance of the animal. A very recent study reported a surprising result that humans who achieved high performance in a planning task tended to replay actions they found to be sub-optimal, and that this was associated with a useful deprecation of those actions in subsequent performance. In this study, we examine the nature of this pessimized form of replay and show that it is exactly appropriate for forgetful agents. We analyse the role of forgetting for replay choices of our model, and verify our predictions using human subject data.
DOI: 10.1016/j.neuron.2011.02.027
发表时间: 2011-03-24
期刊: Neuron
影响因子: 16.2
作者:
Daw ND;Gershman SJ;Seymour B;Dayan P;Dolan RJ
通讯作者: Dolan RJ
DOI: 10.1038/nn.2384
发表时间: 2009-10-01
影响因子: 25
作者:
Girardeau, Gabrielle;Benchenane, Karim;Zugaro, Michael B.
通讯作者: Zugaro, Michael B.
在海马锋利的波浪纹波事件中,协调的激发和抑制前额叶的集合。
DOI: 10.1016/j.neuron.2016.02.010
发表时间: 2016-04-06
期刊: Neuron
影响因子: 16.2
作者:
Jadhav SP;Rothschild G;Roumis DK;Frank LM
通讯作者: Frank LM
DOI: 10.1002/hipo.20707
发表时间: 2010-01
期刊: HIPPOCAMPUS
影响因子: 3.5
作者:
Ego-Stengel, Valerie;Wilson, Matthew A.
通讯作者: Wilson, Matthew A.
DOI: 10.1038/nn1202
发表时间: 2004-03-01
影响因子: 25
作者:
Káli, S;Dayan, P
通讯作者: Dayan, P