课题基金 / 基金详情

A theoretical framework for probabilistic reinforcement learning in the basal ganglia

A theoretical framework for probabilistic reinforcement learning in the basal ganglia
基底神经节概率强化学习的理论框架
批准号:
10460155
负责人:
Samuel J Gershman
金额:
$53.39万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-08-15 至 2024-07-31

项目摘要

项目成果

Samuel J Gershman的其他基金

相似基金

相关文献

中文摘要
翻译
项目摘要 根据标准强化学习框架,基底节实现长时程估计。 长期未来奖励和控制行为,以最大化未来奖励。多巴胺(DA)通过以下途径发挥核心作用 提供指导奖励预测的更新的学习信号(奖励预测误差,或RPE) 行动政策。尽管强化学习框架取得了成功,但它受到了来自 方向数。一些研究表明,DA本身对奖励预测进行编码,而不是 比奖励预测错误,以及其他研究表明,DA可能在激励行动中发挥作用 选择独立于它对学习的贡献。这个项目的一个主要目标是开发一种 基底节功能的强化学习理论解决了这些挑战,更广泛地 提供学习、概率推理和动作选择如何协同工作的统一视图 适应性行为。我们的理论创新可以分为三个部分。首先,我们认为 纹状体的皮质输入编码了隐藏状态的概率分布,称为信念状态。 其次,我们认为纹状体投射神经元通过一组基函数来转换这种输入,其基函数 目的是为了便于奖励预测。将这些预测参数化的突触权重被更新 基于DA RPE信号。第三,我们认为动作选择回路在背侧纹状体使用 关于实施不确定性引导勘探的奖励的概率信息。
英文摘要
Project abstract According to the standard reinforcement learning framework, the basal ganglia implements estimation of long- term future reward and the control of actions to maximize future reward. Dopamine (DA) plays a central role by providing the learning signal (reward prediction error, or RPE) that guides updating of reward predictions and the action policy. Despite its success, the reinforcement learning framework has been challenged from a number of directions. Some studies have suggested that DA encodes reward predictions themselves, rather than reward prediction errors, and other studies have suggested that DA may play a role in invigorating action selection independently from its contribution to learning. A major goal of this project is to develop a reinforcement learning theory of basal ganglia function that addresses these challenges, and more broadly presents a unifying view of how learning, probabilistic inference, and action selection work together to produce adaptive behavior. Our theoretical innovation can be divided into three components. First, we argue that cortical inputs to the striatum encode a probability distribution over hidden states, known as the belief state. Second, we argue that striatal projection neurons transform this input through a set of basis functions, whose purpose is to facilitate reward prediction. The synaptic weights that parametrize these predictions are updated based on the DA RPE signal. Third, we argue that action selection circuits in the dorsal striatum use probabilistic information about rewards to implement uncertainty-guided exploration.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A theoretical framework for probabilistic reinforcement learning in the basal ganglia
  • 批准号:
    10226986
  • 项目类别:
  • 资助金额:
    $52.13万
  • 财政年份:
    2019
  • 负责人:
    Samuel J Gershman
  • 依托单位:
A theoretical framework for probabilistic reinforcement learning in the basal ganglia
  • 批准号:
    10687830
  • 项目类别:
  • 资助金额:
    $53.56万
  • 财政年份:
    2019
  • 负责人:
    Samuel J Gershman
  • 依托单位:
海外基金