Optimism and pessimism in optimised replay.
Optimism and pessimism in optimised replay.
复制标题
DOI:
10.1371/journal.pcbi.1009634
复制
发表时间:
2022-01
影响因子:
4.3
通讯作者:
Dayan P
中科院分区:
文献类型:
--
作者:
Antonov G;Gagne C;Eldar E;Dayan P
The replay of task-relevant trajectories is known to contribute to memory consolidation and improved task performance. A wide variety of experimental data show that the content of replayed sequences is highly specific and can be modulated by reward as well as other prominent task variables. However, the rules governing the choice of sequences to be replayed still remain poorly understood. One recent theoretical suggestion is that the prioritization of replay experiences in decision-making problems is based on their effect on the choice of action. We show that this implies that subjects should replay sub-optimal actions that they dysfunctionally choose rather than optimal ones, when, by being forgetful, they experience large amounts of uncertainty in their internal models of the world. We use this to account for recent experimental data demonstrating exactly pessimal replay, fitting model parameters to the individual subjects’ choices. When animals are asleep or restfully awake, populations of neurons in their brains recapitulate activity associated with extended behaviourally-relevant experiences. This process is called replay, and it has been established for a long time in rodents, and very recently in humans, to be important for good performance in decision-making tasks. The specific experiences which are replayed during those epochs follow highly ordered patterns, but the mechanisms which establish their priority are still not fully understood. One promising theoretical suggestion is that each replay experience is chosen in such a way that the learning that ensues is most helpful for the subsequent performance of the animal. A very recent study reported a surprising result that humans who achieved high performance in a planning task tended to replay actions they found to be sub-optimal, and that this was associated with a useful deprecation of those actions in subsequent performance. In this study, we examine the nature of this pessimized form of replay and show that it is exactly appropriate for forgetful agents. We analyse the role of forgetting for replay choices of our model, and verify our predictions using human subject data.
登录
查看更多内容
影响因子:
16.2
作者:
Daw ND;Gershman SJ;Seymour B;Dayan P;Dolan RJ
通讯作者:
Dolan RJ
影响因子:
25
作者:
Girardeau, Gabrielle;Benchenane, Karim;Zugaro, Michael B.
通讯作者:
Zugaro, Michael B.
影响因子:
16.2
作者:
Jadhav SP;Rothschild G;Roumis DK;Frank LM
通讯作者:
Frank LM
影响因子:
3.5
作者:
Ego-Stengel, Valerie;Wilson, Matthew A.
通讯作者:
Wilson, Matthew A.
影响因子:
25
作者:
Káli, S;Dayan, P
通讯作者:
Dayan, P