Memory-Based Operant Learning
Memory-Based Operant Learning
批准号:
9978403
负责人:
David Touretzky
金额:
$33.83万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1999
资助国家:
美国
项目状态:
已结题
起止时间:
1999-12-15 至 2004-11-30
中文摘要
PI将开发一种认知上合理的强化学习(RL)架构,作为动物和机器人工具学习的模型。虽然强化学习最初是受到动物学习现象的启发,但该领域的发展主要是通过解决人工智能问题。当前RL架构作为认知理论的一个主要限制是状态空间的表示。维持显式状态表示的模型(如Q~表)仅限于只有几个变量的简单域,而隐式表示状态空间的模型(例如,使用神经网络函数逼近器)需要大量的训练数据和与真实动物相比不合理的长时间训练。PI的方法是开发适合动物行为建模的状态空间的专门表示,并且可以支持所需的泛化。模拟动物的工作记忆将编码感官刺激、状态变化事件和动物自己的行为。一个显式的状态表示将对所有这些变量的结合进行编码,从而产生组合爆炸。建议的替代方法是让模型形成选定变量的连词,允许它在只关注与正在学习的任务相关的那些维度时,增量地扩展其状态描述。将开发基于快速单层神经网络学习的启发式方法,以根据最近的经验选择有用的连词。PI还将研究将工作记忆的当前状态与整个过去状态或事件的记录相匹配,以预测奖励。将开发一种灵活的架构,用于以参数化的方式表示动作(以便提供无限的可变性),并具有时间持续时间(允许在执行过程中到达刺激和奖励)。最后,将存在一些机制来应对行动的失败,以成功执行或产生预期奖励;这将为建模现象提供基础,如部分强化时间表的影响和灭绝期间增加的行为可变性。如果成功,这项工作将通过引入处理复杂状态和动作空间的新技术来推进强化学习的发展。这对动物认知理论、通过探索和实验学习的机器人以及向人类老师学习的机器人都具有重要意义。
英文摘要
The PI will develop a cognitively plausible reinforcement learning (RL) architecture as a model of instrumental learning in animals and robots. Although RL was initially inspired by animal learning phenomena, the field has since developed mainly by addressing AI concerns. A major limitation of current RL architectures as cognitive theories is the representation of state space. Models that maintain explicit state representations (such as Q~tables) are limited to simple domains with only a few variables, while models that represent the state space implicitly (e.g., using a neural net function approximator) require large amounts of training data and unreasonably long training times compared to real animals. The PI's approach is to develop specialized representations of state space that are appropriate for modeling animal behavior and can support desired generalizations. The simulated animal's working memory will encode sensory stimuli, state change events, and the animal's own actions. An explicit state representation would encode the conjunction of all these variables, generating a combinatorial explosion. The proposed alternative approach is for the model to form conjunctions of selected variables, allowing it to incrementally expand its state description while focusing on just those dimensions that are relevant to the task being learned. Heuristics based on fast, single-layer neural net learning will be developed to select useful conjunctions as a function of recent experience. The PI also will investigate matching the current state of working memory with records of entire past states, or episodes, in order to predict reward. A flexible architecture will be developed for representing actions in a parameterized manner (so as to provide infinite variability), and with temporal duration (allowing stimuli and rewards to arrive in the midst of execution). Finally, there will be mechanisms for coping with failure of an action to execute successfully or to produce an expected reward; this will provide the basis for modeling phenomena such as effects of partial reinforcement schedules and increased behavioral variability during extinction. If successful, this work will advance the state of the art of reinforcement learning by introducing new techniques for handling complex state and action spaces. This has important implications for theories of animal cognition, for robots that learn by exploration and experimentation, and for robots intended to learn from human teachers.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: AI4GA - Developing Artificial Intelligence Competencies, Career Awareness, and Interest in Georgia Middle School Teachers and Students
-
批准号:2049029
-
项目类别:Standard Grant
-
资助金额:$101.68万
-
财政年份:2021
-
负责人:David Touretzky
-
依托单位:
Developing K-12 Education Guidelines for Artificial Intelligence
-
批准号:1846073
-
项目类别:Standard Grant
-
资助金额:$22.57万
-
财政年份:2019
-
负责人:David Touretzky
-
依托单位:
Collaborative Research: Planning grant: CS4All: Computer Science for All
-
批准号:1151542
-
项目类别:Standard Grant
-
资助金额:$4.26万
-
财政年份:2012
-
负责人:David Touretzky
-
依托单位:
BPC-AE: Collaborative Research: The ARTSI Alliance: Advancing Robotics Technology for Societal Impact
-
批准号:1042322
-
项目类别:Continuing Grant
-
资助金额:$20.0万
-
财政年份:2011
-
负责人:David Touretzky
-
依托单位:
Cognitive Robotics: A Curriculum for Machines that See and Manipulate their World
-
批准号:0717705
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:David Touretzky
-
依托单位:
Collaborative Research: BPC-A: ARTSI: Advancing Robotics Technology for Societal Impact
-
批准号:0742106
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:David Touretzky
-
依托单位:
Collaborative Research: BPC-DP: CARE. Computer and Robotics Education for African American Students
-
批准号:0540521
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2006
-
负责人:David Touretzky
-
依托单位:
IGERT: Integrating New Technologies with Cognitive Neuroscience
-
批准号:0549352
-
项目类别:Continuing Grant
-
资助金额:$254.0万
-
财政年份:2006
-
负责人:David Touretzky
-
依托单位:
IGERT Full Proposal: Innovative Cross-Disciplinary Training in Neuroscience and Computation
-
批准号:9987588
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2000
-
负责人:David Touretzky
-
依托单位:
The Biological Basis of Incremental Learning
-
批准号:9720350
-
项目类别:Continuing Grant
-
资助金额:$77.5万
-
财政年份:1997
-
负责人:David Touretzky
-
依托单位:
A Computational Theory of Operant Conditioning with Application to Trainable Robots
-
批准号:9530975
-
项目类别:Continuing Grant
-
资助金额:$25.66万
-
财政年份:1996
-
负责人:David Touretzky
-
依托单位:
Computational Modeling of the Rodent Head Direction System
-
批准号:9631336
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:1996
-
负责人:David Touretzky
-
依托单位:
Inheritance Theory and Knowledge Bases
-
批准号:9003165
-
项目类别:Continuing Grant
-
资助金额:$27.08万
-
财政年份:1990
-
负责人:David Touretzky
-
依托单位:
Distributed Representations for Symbolic Data Structures (Information Science)
-
批准号:8516330
-
项目类别:Standard Grant
-
资助金额:$6.12万
-
财政年份:1986
-
负责人:David Touretzky
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
-
批准号:--
-
项目类别:外国青年学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:江洋子
-
依托单位:
Incentive and governance schenism study of corporate green washing behavior in China: Based on an integiated view of econfiguration of environmental authority and decoupling logic
-
批准号:--
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:YU BYUNGJUN
-
依托单位:
Exploring the Intrinsic Mechanisms of CEO Turnover and Market Reaction: An Explanation Based on Information Asymmetry
-
批准号:W2433169
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:HAOFEI ZHANG
-
依托单位:
A study on prototype flexible multifunctional graphene foam-based sensing grid (柔性多功能石墨烯泡沫传感网格原型研究)
-
批准号:--
-
项目类别:--
-
资助金额:20万元
-
批准年份:2020
-
负责人:SAGAR RIZWAN UR REHMAN
-
依托单位:
基于tag-based单细胞转录组测序解析造血干细胞发育的可变剪接
-
批准号:81900115
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:李宗城
-
依托单位:
应用Agent-Based-Model研究围术期单剂量地塞米松对手术切口愈合的影响及机制
-
批准号:81771933
-
项目类别:面上项目
-
资助金额:50.0万元
-
批准年份:2017
-
负责人:周全红
-
依托单位:
Reality-based Interaction用户界面模型和评估方法研究
-
批准号:61170182
-
项目类别:面上项目
-
资助金额:57.0万元
-
批准年份:2011
-
负责人:田丰
-
依托单位:
Multistage,haplotype and functional tests-based FCAR 基因和IgA肾病相关关系研究
-
批准号:30771013
-
项目类别:面上项目
-
资助金额:30.0万元
-
批准年份:2007
-
负责人:王一鸣
-
依托单位:
差异蛋白质组技术结合Array-based CGH 寻找骨肉瘤分子标志物
-
批准号:30470665
-
项目类别:面上项目
-
资助金额:8.0万元
-
批准年份:2004
-
负责人:李扬
-
依托单位:
GaN-based稀磁半导体材料与自旋电子共振隧穿器件的研究
-
批准号:60376005
-
项目类别:面上项目
-
资助金额:20.0万元
-
批准年份:2003
-
负责人:张国义
-
依托单位: