Circuit mechanisms of arbitration between distinct reinforcement learning systems
Circuit mechanisms of arbitration between distinct reinforcement learning systems
批准号:
10608739
负责人:
Margaret Louise DeMaegd
金额:
$7.41万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-03-01 至 2026-02-28
关键词:
AddressAnatomyAnimalsArbitrationBehaviorBehavior monitoringBehavioral ModelBrainChronicComputing MethodologiesConfocal MicroscopyCorpus striatum structureDataDecision MakingDiseaseDorsalElectrophysiology (science)EnvironmentExcitatory Postsynaptic PotentialsExhibitsFutureGeneticGoalsHumanImplantIn VitroInjectionsInterneuronsKnowledgeLateralLeadLearningLogicMeasuresMedialMediatingMethodsMicroelectrodesModelingNeuronsObsessive-Compulsive DisorderPatternPopulationPrefrontal CortexProcessPsyche structurePsychological reinforcementRattusReportingStructureSynapsesSystemTechniquesTestingTheoretical StudiesTracerUncertaintyViralVisualWhole-Cell RecordingsWorkbehavioral studycell typecognitive taskexperienceflexibilityfree behaviorin vivomicroscopic imagingneuralneural circuitnovelnovel therapeuticsoptogeneticsresponsesimulationstatistics
中文摘要
项目摘要
动物可以在新的环境中表现出目标导向的行为,尽管经验有限
和他们大脑是如何对潜在的统计数据做出和使用推断的,
环境的生成结构来指导行为?强化学习领域指的是
这种能力是“基于模型的”推理,这意味着它依赖于一个内部模型,
世界的结构。重要的是,这种内部模型可以用来灵活地估计最佳的
行动通过心理模拟或计划,没有直接的经验。相反,在“无模型”中,
强化学习,智能体根据直接经验选择最佳行动,而不需要
任务或环境的底层顺序转换结构的明确知识。
基于模型和无模型的机制在大脑中共存,并由不同的神经元介导。
电路,虽然大脑在这些电路之间进行仲裁的神经电路机制
决策系统仍然未知。理论和行为研究表明,
大脑使用的系统产生的价值估计具有最低的不确定性。横向
眶额皮层(IOFC)是执行仲裁的一个令人信服的候选者,因为虽然它是
涉及基于模型的推理,例如通过启用关于隐藏任务的推断
它位于背侧纹状体的上游,这对基于模型和模型的研究都至关重要。
自由决策。有趣的是,我们发现lOFC神经元只投射到
背外侧纹状体(DLS),一个对无模型行为至关重要的区域,而不是背内侧
纹状体(DMS),这对基于模型的行为至关重要。我们假设投射
IOFC中的特定神经回路通过抑制无模型信号在这些系统之间进行仲裁,
系统
我将使用最先进的病毒、电生理和计算方法来
确定DLS投射的IOFC神经元是否介导DLS投射的IOFC神经元之间的基于不确定性的仲裁。
决策系统(目标1),并表征支持的底层电路逻辑
仲裁(目标2)。通过光遗传学标记DLS投射的IOFC神经元,
在监测大鼠在任务中使用的行为策略时,
潜在的结构。为了确定仲裁是如何在背侧纹状体中实例化的,我将
光遗传学激活OFC→DLS神经元,同时记录来自不同遗传细胞类型的
纹状体,在体内和体外。我们预测OFC→DLS神经元能够实现基于模型的
通过激活抑制性中间神经元来抑制DLS和无模型系统的行为。
英文摘要
PROJECT SUMMARY
Animals can exhibit goal-directed behaviors in novel environments, despite limited experience
with them. How does the brain make and use inferences about the underlying statistics and
generative structure of environments to guide behavior? The field of reinforcement learning refers
to this capacity as “model-based” reasoning, meaning that it relies on an internal model of the
structure of the world. Critically, this internal model can be used to flexibly estimate the best
actions by mental simulation or planning, without direct experience. In contrast, in “model-free”
reinforcement learning, an agent chooses the best action based on direct experience, without
explicit knowledge of the underlying sequential transition structure of a task or environment.
Model-based and model-free mechanisms coexist in the brain and are mediated by distinct
circuits, although the neural circuit mechanisms by which the brain arbitrates between these
decision systems remains unknown. Theoretical and behavioral studies suggest that human
brains use the system that yields value estimates with the lowest uncertainty. The lateral
orbitofrontal cortex (lOFC) is a compelling candidate to perform arbitration because while it is
implicated in model-based reasoning, for instance by enabling inferences about hidden task
states, it lies upstream of the dorsal striatum, which is critical for both model-based and model-
free decision making. Intriguingly, we have found that lOFC neurons project exclusively to the
dorsolateral striatum (DLS), a region critical for model-free behavior, and not the dorsomedial
striatum (DMS), which is critical for model-based behavior. We hypothesize that projection
specific neural circuits in lOFC arbitrate between these systems by suppressing the model-free
system.
I will use state-of-the-art viral, electrophysiological, and computational methods to
determine whether DLS-projecting lOFC neurons mediate uncertainty-based arbitration between
decision-making systems (Aim 1) and characterize the underlying circuit logic that supports
arbitration (Aim 2). By optogenetically tagging DLS-projecting lOFC neurons I will selectively
characterize and perturb their activity while monitoring the behavioral strategy rats use in a task
with latent structure. To determine how arbitration is instantiated in the dorsal striatum I will
optogenetically activate OFC→DLS neurons while recording from different genetic cell types in
the striatum, in vivo and in vitro. We predict that OFC→DLS neurons enable model-based
behavior by activating inhibitory interneurons to suppress the DLS and the model-free system.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金