Dynamic Abstraction in Reinforcement Learning
Dynamic Abstraction in Reinforcement Learning
批准号:
0218125
负责人:
Andrew Barto
金额:
$19.96万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2002
资助国家:
美国
项目状态:
已结题
起止时间:
2002-09-01 至 2005-08-31
中文摘要
该项目研究了使用动态抽象来利用复杂环境的时空结构来促进学习的强化学习算法。抽象的使用是人类智能的特征之一,它使我们能够像在复杂的环境中一样有效地运作。我们系统地忽略了与手头任务无关的细节,当我们专注于一系列子任务时,我们会在抽象之间快速切换。例如,在计划日常活动时,比如开车去上班,我们会抽象出不相关的细节,比如车内物体的布局,但当我们真正开车时,很多细节就变得相关了,比如方向盘和油门的位置。不同的抽象适用于不同的任务或子任务,代理必须在转移到新任务或新子任务时转移抽象。这个项目结合了选项理论与因子状态和动作表示,以赋予动态抽象概念精确的含义,并研究创建和利用它们的方法。它将通过将现有的单步动态贝叶斯网络模型的形式化扩展到多时间情况,开发用于在因子状态和动作表示方面表示期权模型的形式化。它将研究多次公式调用如何促进动态抽象的创建和使用。通过将经典自动机理论的相关概念扩展到多时间因子模型,将发展一种代数抽象理论。通过扩展现有的混合模型算法,从单步学习过渡模型到多步模型,将开发用于学习紧凑多步选择模型的方法。一般来说,动态抽象的概念将是一个有价值的工具,适用于大规模制造(例如,工厂过程控制),机器人(导航),多代理协调和其他最先进的强化学习应用中的许多困难的优化问题。由于这项研究结合了决策理论、运筹学、控制理论、认知科学和人工智能领域的思想,它可能提供一个有用的桥梁,有可能促进所有这些领域的贡献。
英文摘要
This project investigates reinforcement learning algorithms that use dynamic abstraction to exploit the spatial and temporal structure of complex environments to facilitate learning. The use of abstraction is one of the features of human intelligence that allows us to operate as effectively as we do in complex environments. We systematically ignore details that are not relevant to a task at hand, and we rapidly switch between abstractions when we focus on a succession of subtasks. For example, in planning everyday activities, such as driving to work, we abstract out irrelevant details such as the layout of objects inside the car, but when we actually drive, many of these details become relevant, such as the locations of the steering wheel and the accelerator. Different abstractions are appropriate for different tasks or subtasks, and the agent has to shift abstractions as it shifts to new tasks or to new subtasks.This project combines the theory of options with factored state and action representations to give precise meaning to the concept of dynamic abstractions and to study methods for creating and exploiting them. It will develop formalisms for representing option models in terms of factored state and action representations by extending existing formalisms for single-step dynamic Bayes network models to the multi-time case. It will investigate how the multi-time formulation call facilitate creating and using dynamic abstractions. An algebraic theory of abstraction will be developed by extending relevant concepts from classical automata theory to multi-time factored models. Methods will be developed for learning compact multistep option models by extending an existing mixture model algorithm for learning transition models from single-step to multi-step models. In general the notion of dynamic abstraction will be a valuable tool to apply to many difficult optimization problems in large-scale manufacturing (e.g., factory process control), robotics (navigation), multi-agent coordination, and other state-of-the-art applications of reinforcement learning. Since this research combines ideas from the fields of decision theory, operations research, control theory, cognitive science, and AI, it may provide a useful bridge that has the potential to foster contributions in all of these fields.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CRCNS: Collaborative Research: Neural Correlates of Hierarchical Reinforcement Learning
-
批准号:1208051
-
项目类别:Continuing Grant
-
资助金额:$4.33万
-
财政年份:2012
-
负责人:Andrew Barto
-
依托单位:
NRI-Small: Collaborative Research: Multiple Task Learning from Unstructured Demonstrations
-
批准号:1208497
-
项目类别:Standard Grant
-
资助金额:$49.92万
-
财政年份:2012
-
负责人:Andrew Barto
-
依托单位:
SGER: Building Blocks for Creative Search
-
批准号:0733581
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:Andrew Barto
-
依托单位:
Collaborative Research: Intrinsically Motivated Learning in Artificial Agents
-
批准号:0432143
-
项目类别:Continuing Grant
-
资助金额:$15.0万
-
财政年份:2004
-
负责人:Andrew Barto
-
依托单位:
Lyapunov Methods for Reinforcement Learning
-
批准号:0070102
-
项目类别:Standard Grant
-
资助金额:$11.94万
-
财政年份:2000
-
负责人:Andrew Barto
-
依托单位:
KDI: Temporal Abstraction in Reinforcement Learning
-
批准号:9980062
-
项目类别:Standard Grant
-
资助金额:$56.0万
-
财政年份:1999
-
负责人:Andrew Barto
-
依托单位:
Multiple Time Scale Reinforcement Learning
-
批准号:9511805
-
项目类别:Continuing Grant
-
资助金额:$15.73万
-
财政年份:1995
-
负责人:Andrew Barto
-
依托单位:
Reinforcement Learning Algorithms Based on Dynamic Programming
-
批准号:9214866
-
项目类别:Continuing Grant
-
资助金额:$31.63万
-
财政年份:1992
-
负责人:Andrew Barto
-
依托单位:
Neural Networks for Adaptive Control
-
批准号:8912623
-
项目类别:Continuing Grant
-
资助金额:$63.92万
-
财政年份:1989
-
负责人:Andrew Barto
-
依托单位:
Conference on the Neurone as a Computational Unit, June 28--July 1, 1988, King's College, Cambridge, England
-
批准号:8808758
-
项目类别:Standard Grant
-
资助金额:$1.64万
-
财政年份:1988
-
负责人:Andrew Barto
-
依托单位:
US-United Kingdom Cooperative Science: A Theoretical Study of the Perceptual Prerequisites of Learning
-
批准号:8815252
-
项目类别:Standard Grant
-
资助金额:$1.14万
-
财政年份:1988
-
负责人:Andrew Barto
-
依托单位:
海外基金