Decision dynamics during a continuous-time foraging task: a reinforcement learning approach
Decision dynamics during a continuous-time foraging task: a reinforcement learning approach
批准号:
10373999
负责人:
Benjamin Nicolaas Ballintyn
金额:
$0.52万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-04-01 至 2022-04-15
关键词:
AlgorithmsAnimal BehaviorAnimalsBehaviorBehavioralBehavioral ModelBrainBreathingComplementComputer ModelsConsumptionDataDecision MakingDecision ModelingDevelopmentEconomicsElectrophysiology (science)Energy IntakeEnvironmentEvolutionFamilyFoodFree WillFutureGoalsHourHumanIndividualKnowledgeLearningLearning DisordersLifeLocationMeasurementMeasuresMemoryMethodsModelingNatural SelectionsNaturePalatePathologyPerformancePoliciesProbabilityProcessPsychological reinforcementRattusResearchResourcesRewardsSamplingSelf-control as a personality traitSourceStimulusStructureSystemTestingThirstTimeTravelUpdateWait TimeWeightaddictionbasebehavior predictioncognitive processdecision making algorithmdiscountingexperimental studyimprovedindexinginsightlearning algorithmlearning strategymotor disorderneural circuitneurophysiologypressurerelating to nervous systemreward circuitrysuccesstheories
中文摘要
项目概要
进化很可能强烈塑造了奖励系统的神经回路以优化
执行与寻找资源有关的许多任务,这是每只动物生命的重要组成部分。这个
这个命题是“最优”觅食理论发展的灵感来源,例如边际
价值定理(MVT),通过分析推导出觅食行为(选择序列)
最大化长期奖励率,通常被认为是能量摄入。虽然这些分析
理论在描述动物行为方面取得了一些成功,理论本身依赖于严格的
关于环境的假设在许多自然情况下并不成立,而且不够灵活
推广到更复杂的环境或其他任务。因此,该项目的最终目标是
了解通用决策(强化学习)算法系列中的哪一个最有可能
被大脑用来解决基于价值的任务,并利用这些知识来预测什么
在某些情况下,这些算法将导致最佳或次优的行为。
通过这个项目,我将提高我们对这些更自然的动物决策过程的理解
通过执行时间上连续且违反许多规则的觅食实验
先前的觅食分析理论所依赖的假设。因口渴而引起的老鼠将被允许
在开放场地中从两个或三个(可口的或令人厌恶的)促味剂选项(“补丁”)中自由取样,并且,
关键的是,将被允许指导他们与选项的接触,这是过去的实验所经历过的
缺乏。对每个促味剂选项的舔(消费)行为的测量将使我能够
测量大鼠在几次 1 小时的会议中的决策动态。特别是,我将衡量如何
每个选项的采样时间与替代方案的值相关,以深入了解老鼠如何
结合可用选项的值来做出决策。
作为此行为任务的补充,我将模拟一组强化学习代理,这些代理的不同之处在于
用于学习行动价值观、选择行动和规划行动的规则。通过定量
将这些人工代理的决策行为与从老鼠身上获得的决策行为进行比较,我将确定哪一个
模拟代理最好地再现了老鼠的行为,从而深入了解了所使用的决策算法
大鼠并为这项任务中未来的电生理记录提供方向。重要的是,这
将动物行为与人工行为进行比较将使我能够评估与人工行为的接近程度
“最佳”大鼠行为是,并且在次优的情况下,为以下行为提供定量解释:
为什么会这样。
英文摘要
Project Summary
It is likely that evolution has strongly shaped the neural circuitry of the reward systems to optimize
performance in the many tasks involved in foraging for resources, a critical part of every animal's life. This
proposition was the inspiration for the development of “optimal” foraging theories, such as the marginal
value theorem (MVT), which derive analytically the foraging behavior (sequences of choices) that
maximizes the long-term rate of reward, usually considered to be energy intake. While these analytic
theories have had some success in describing animal behavior, the theories themselves rely on strict
assumptions about the environment that do not hold in many natural situations and are not flexible enough
to generalize to more complicated environments or other tasks. Therefore, the end goal of this project is to
understand which of a family of general-purpose decision (reinforcement-learning) algorithms is most likely
to be employed by the brain to solve value-based tasks and to use this knowledge to predict under what
circumstances these algorithms will lead to optimal or suboptimal behavior.
With this project, I will improve our understanding of animal decision processes in these more natural
environments by performing a foraging experiment that is continuous in time and violates many of the
assumptions that prior analytical theories of foraging rely on. Rats motivated by thirst will be allowed to
sample freely from two or three (palatable or aversive) tastant options (“patches”) in an open field and,
critically, will be allowed to direct their encounters with the options, something which past experiments have
lacked. Measurements of licking (consumption) behavior at each of the tastant options will allow me to
measure the decision dynamics of the rat over several 1-hour sessions. In particular, I will measure how the
sampling times at each option correlate with the values of the alternatives to gain insight into how rats
combine the values of available options to make decisions.
As a complement to this behavioral task, I will simulate a set of reinforcement learning agents that vary in
the rules used for learning action values, choosing actions, and planning actions. By quantitatively
comparing the decision behavior of these artificial agents to that obtained from rats I will determine which of
the simulated agents best reproduces the rat behavior, giving insight into the decision algorithms used by
rats and providing a direction for future electrophysiological recordings during this task. Importantly, this
comparison of animal behavior with that produced by artificial agents will allow me to assess how close to
“optimal” rat behavior is and, in the cases where it is suboptimal, to provide quantitative explanations for
why it is so.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金