Few-shot learning: temporal scaling in behavioral and dopaminergic learning.

Few-shot learning: temporal scaling in behavioral and dopaminergic learning.
复制标题

少样本学习:行为和多巴胺能学习中的时间尺度。

DOI:
10.1101/2023.03.31.535173
复制
发表时间:
2023
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
Namboodiri,VijayMohanK
Namboodiri,VijayMohanK
中科院分区:
--
文献类型:
--
作者:
Burke,DennisA;Jeong,Huijeong;Wu,Brenda;Lee,SeulAh;Floeder,JosephR;Namboodiri,VijayMohanK

文献摘要

相似文献

我们如何学习世界中的关联(例如,提示和奖励之间的关联)?大脑中的提示奖励联想学习是由中脑边缘多巴胺控制的。人们普遍认为,多巴胺通过根据时间差强化学习(TDRL)算法传递奖励预测误差(RPE)来驱动这种学习。 TDRL 的实施是“基于试验的”:学习在个人提示结果体验中按顺序进行。因此,一个基本假设(通常被认为是不言而喻的事实)是,一个人经历的提示奖励配对越多,就越能学到这种关联。在这里,我们反驳了这个假设,从而伪造了基于试验的学习算法的基本原理。具体来说,当一组固定头部的小鼠在相同的总时间内获得的经验是另一组小鼠的十倍时,一次经历产生的学习量与另一组的十次经历产生的学习量一样多。这种定量尺度也适用于中脑边缘多巴胺能学习,学习率的增加是如此之高,以至于经验较少的组在短短四次提示奖励体验中表现出多巴胺能学习,在九次中表现出行为学习。实施奖励触发回顾性学习的算法解释了这些发现。这里观察到的时间缩放和小样本学习从根本上改变了我们对联想学习神经算法的理解。
How do we learn associations in the world (e.g., between cues and rewards)? Cue-reward associative learning is controlled in the brain by mesolimbic dopamine–. It is widely believed that dopamine drives such learning by conveying a reward prediction error (RPE) in accordance with temporal difference reinforcement learning (TDRL) algorithms. TDRL implementations are “trial-based”: learning progresses sequentially across individual cue-outcome experiences. Accordingly, a foundational assumption—often considered a mere truism—is that the more cuereward pairings one experiences, the more one learns this association. Here, we disprove this assumption, thereby falsifying a foundational principle of trial-based learning algorithms. Specifically, when a group of head-fixed mice received ten times fewer experiences over the same total time as another, a single experience produced as much learning as ten experiences in the other group. This quantitative scaling also holds for mesolimbic dopaminergic learning, with the increase in learning rate being so high that the group with fewer experiences exhibits dopaminergic learning in as few as four cue-reward experiences and behavioral learning in nine. An algorithm implementing reward-triggered retrospective learning explains these findings. The temporal scaling and few-shot learning observed here fundamentally changes our understanding of the neural algorithms of associative learning.