Inferring learning rules from animal decision-making

Inferring learning rules from animal decision-making
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
Advances in neural information processing systems
影响因子:
--
通讯作者:
Zoe C. Ashwood;Nicholas A. Roy;Ji Hyun Bak;Jonathan W. Pillow
Zoe C. Ashwood;Nicholas A. Roy;Ji Hyun Bak;Jonathan W. Pillow
中科院分区:
其他
文献类型:
--
作者:
Zoe C. Ashwood;Nicholas A. Roy;Ji Hyun Bak;Jonathan W. Pillow

文献摘要

被引文献

相似文献

动物是如何学习的?这在神经科学中仍然是一个难以捉摸的问题。强化学习通常侧重于算法的设计,使人工智能体能够有效地学习新任务,而在这里,我们开发了一个建模框架来直接推断动物用来获得新行为的经验学习规则。我们的方法有效地推断出动物策略中试验到试验的变化,并将这些变化分解为学习成分和噪声成分。具体来说,这允许我们:(i)比较动物可能使用的不同学习规则和目标函数来更新其策略;(ii)对动物政策的不同参数估计不同的学习率;(iii)确定不同动物群体的学习差异;(iv)发现没有被规范学习规则捕获的试验到试验的变化。在模拟选择数据上验证了我们的框架后,我们将我们的模型应用于学习感知决策任务的大鼠和小鼠的数据。我们发现,某些学习规则更能解释动物政策在一次又一次试验中的变化。而传统强化学习规则对小鼠学习国际脑实验室任务的策略更新的平均贡献仅为30%,我们发现,在我们的模型下,添加基线参数允许学习规则解释92%的动物策略更新。有趣的是,最佳拟合学习率和基线值表明,在每次试验中,动物的策略更新并不会朝着预期奖励最大化的方向发生。了解动物在学习一项新任务时如何从偶然表现过渡到高精度表现,不仅为神经科学家提供了对动物的深入了解,还为机器学习社区提供了生物学习算法的具体示例。
How do animals learn? This remains an elusive question in neuroscience. Whereas reinforcement learning often focuses on the design of algorithms that enable artificial agents to efficiently learn new tasks, here we develop a modeling framework to directly infer the empirical learning rules that animals use to acquire new behaviors. Our method efficiently infers the trial-to-trial changes in an animal's policy, and decomposes those changes into a learning component and a noise component. Specifically, this allows us to: (i) compare different learning rules and objective functions that an animal may be using to update its policy; (ii) estimate distinct learning rates for different parameters of an animal's policy; (iii) identify variations in learning across cohorts of animals; and (iv) uncover trial-to-trial changes that are not captured by normative learning rules. After validating our framework on simulated choice data, we applied our model to data from rats and mice learning perceptual decision-making tasks. We found that certain learning rules were far more capable of explaining trial-to-trial changes in an animal's policy. Whereas the average contribution of the conventional REINFORCE learning rule to the policy update for mice learning the International Brain Laboratory's task was just 30%, we found that adding baseline parameters allowed the learning rule to explain 92% of the animals' policy updates under our model. Intriguingly, the best-fitting learning rates and baseline values indicate that an animal's policy update, at each trial, does not occur in the direction that maximizes expected reward. Understanding how an animal transitions from chance-level to high-accuracy performance when learning a new task not only provides neuroscientists with insight into their animals, but also provides concrete examples of biological learning algorithms to the machine learning community.