课题基金 / 基金详情

Memory-Based Operant Learning

Memory-Based Operant Learning
基于记忆的操作学习
批准号:
9978403
负责人:
David Touretzky
金额:
$33.83万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1999
资助国家:
美国
项目状态:
已结题
起止时间:
1999-12-15 至 2004-11-30

项目摘要

项目成果

David Touretzky的其他基金

相似基金

相关文献

中文摘要
翻译
PI将开发一种认知上合理的强化学习(RL)架构,作为动物和机器人工具性学习的模型。尽管RL最初的灵感来自动物学习现象,但自那以来,该领域的发展主要是通过解决人工智能问题。作为认知理论的当前RL体系结构的一个主要限制是状态空间的表示。与真实动物相比,保持显式状态表示的模型(如Q表)仅限于具有少量变量的简单域,而隐式表示状态空间的模型(例如,使用神经网络函数逼近器)需要大量的训练数据和不合理的长时间训练。PI的方法是开发适合于对动物行为建模的状态空间的专门表示,并且可以支持所需的泛化。模拟动物的工作记忆将编码感觉刺激、状态变化事件和动物自己的行为。显式的状态表示将编码所有这些变量的结合,从而产生组合爆炸。建议的替代方法是让模型形成选定变量的连接,允许它增量地扩展其状态描述,同时只关注与正在学习的任务相关的维度。基于快速单层神经网络学习的启发式将被开发来根据最近的经验来选择有用的连词。PI还将研究将工作记忆的当前状态与整个过去状态或事件的记录进行匹配,以便预测奖励。将开发一种灵活的体系结构,以参数方式(以便提供无限的可变性)和时间持续时间(允许刺激和奖励在执行过程中到达)来表示行动。最后,将有机制来处理未能成功执行或产生预期奖励的行动;这将为模拟部分增援计划的影响和灭绝期间行为变异性增加等现象提供基础。如果成功,这项工作将通过引入处理复杂状态和动作空间的新技术来推动强化学习的技术水平。这对动物认知理论、通过探索和实验学习的机器人以及打算向人类老师学习的机器人都有重要的影响。
英文摘要
The PI will develop a cognitively plausible reinforcement learning (RL) architecture as a model of instrumental learning in animals and robots. Although RL was initially inspired by animal learning phenomena, the field has since developed mainly by addressing AI concerns. A major limitation of current RL architectures as cognitive theories is the representation of state space. Models that maintain explicit state representations (such as Q~tables) are limited to simple domains with only a few variables, while models that represent the state space implicitly (e.g., using a neural net function approximator) require large amounts of training data and unreasonably long training times compared to real animals. The PI's approach is to develop specialized representations of state space that are appropriate for modeling animal behavior and can support desired generalizations. The simulated animal's working memory will encode sensory stimuli, state change events, and the animal's own actions. An explicit state representation would encode the conjunction of all these variables, generating a combinatorial explosion. The proposed alternative approach is for the model to form conjunctions of selected variables, allowing it to incrementally expand its state description while focusing on just those dimensions that are relevant to the task being learned. Heuristics based on fast, single-layer neural net learning will be developed to select useful conjunctions as a function of recent experience. The PI also will investigate matching the current state of working memory with records of entire past states, or episodes, in order to predict reward. A flexible architecture will be developed for representing actions in a parameterized manner (so as to provide infinite variability), and with temporal duration (allowing stimuli and rewards to arrive in the midst of execution). Finally, there will be mechanisms for coping with failure of an action to execute successfully or to produce an expected reward; this will provide the basis for modeling phenomena such as effects of partial reinforcement schedules and increased behavioral variability during extinction. If successful, this work will advance the state of the art of reinforcement learning by introducing new techniques for handling complex state and action spaces. This has important implications for theories of animal cognition, for robots that learn by exploration and experimentation, and for robots intended to learn from human teachers.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: AI4GA - Developing Artificial Intelligence Competencies, Career Awareness, and Interest in Georgia Middle School Teachers and Students
  • 批准号:
    2049029
  • 项目类别:
    Standard Grant
  • 资助金额:
    $101.68万
  • 财政年份:
    2021
  • 负责人:
    David Touretzky
  • 依托单位:
Developing K-12 Education Guidelines for Artificial Intelligence
  • 批准号:
    1846073
  • 项目类别:
    Standard Grant
  • 资助金额:
    $22.57万
  • 财政年份:
    2019
  • 负责人:
    David Touretzky
  • 依托单位:
Collaborative Research: Planning grant: CS4All: Computer Science for All
  • 批准号:
    1151542
  • 项目类别:
    Standard Grant
  • 资助金额:
    $4.26万
  • 财政年份:
    2012
  • 负责人:
    David Touretzky
  • 依托单位:
BPC-AE: Collaborative Research: The ARTSI Alliance: Advancing Robotics Technology for Societal Impact
  • 批准号:
    1042322
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $20.0万
  • 财政年份:
    2011
  • 负责人:
    David Touretzky
  • 依托单位:
国内基金
海外基金
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
Incentive and governance schenism study of corporate green washing behavior in China: Based on an integiated view of econfiguration of environmental authority and decoupling logic
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    YU BYUNGJUN
  • 依托单位:
Exploring the Intrinsic Mechanisms of CEO Turnover and Market Reaction: An Explanation Based on Information Asymmetry
  • 批准号:
    W2433169
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    HAOFEI ZHANG
  • 依托单位:
A study on prototype flexible multifunctional graphene foam-based sensing grid (柔性多功能石墨烯泡沫传感网格原型研究)
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    20万元
  • 批准年份:
    2020
  • 负责人:
    SAGAR RIZWAN UR REHMAN
  • 依托单位: