课题基金 / 基金详情

KDI: Temporal Abstraction in Reinforcement Learning

KDI: Temporal Abstraction in Reinforcement Learning
KDI:强化学习中的时间抽象
批准号:
9980062
负责人:
Andrew Barto
金额:
$56.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
1999
资助国家:
美国
项目状态:
已结题
起止时间:
1999-09-15 至 2003-08-31

项目摘要

项目成果

Andrew Barto的其他基金

相似基金

相关文献

中文摘要
翻译
该项目研究了一种新的方法来学习,规划和表示知识的多层次的时间抽象。 它开发了人工强化学习系统可以对持续的行动过程进行建模和推理并以相应的方式感知其环境的方法,并开发和检查与这种方法相关的动物行为模型的有效性。 该项目的目标是发展该方法的数学理论,完善,扩展和进行动物行为相关模型的验证研究,研究该理论与控制理论和人工智能的关系,并在一些模拟学习任务中证明其有效性。目前最流行的强化学习(RL)研究使用了一个框架,在这个框架中,智能体必须采取一系列由单个固定时间步长控制的动作:动作采取一个步骤来完成,并且它们的直接后果在一个步骤之后变得可用(建模为马尔可夫决策过程,或MDP)。 这使得在不同的时间尺度上学习和计划变得困难。一些强化学习研究使用了这种框架的推广(半马尔可夫决策过程,或SMDP),其中动作需要不同的时间来完成,现有的理论指定了如何对这些动作的结果进行建模以及如何使用它们进行规划。 然而,这种方法是有限的,因为时间上扩展的行动被视为不可分割的和未知的单位。 为了获得最大的灵活性和最佳性能,有必要在时间上扩展的动作中检查或修改它们是如何由较低级别的动作组成的,这在现有的方法中是没有考虑的。相比之下,这个项目将扩展的动作过程建模为SMDP动作覆盖在基础MDP上。 这些行动过程,称为选项,可以被视为原始行动;现有的强化学习几乎可以不加改变地应用。 这种方法使选项能够在MDP和SMDP两个级别进行分析,并在两个级别之间的接口引入新的问题。 这种方法很有吸引力,因为它的简单性,与以前使用原始动作的方法相似,以及它在MDP和SMDP理论中的坚实数学基础。 这正在进一步发展成为一种分级和多时间尺度规划和学习的一般方法。
英文摘要
This project investigates a new approach to learning, planning, and representing knowledge at multiple levels of temporal abstraction. It develops methods by which an artificial reinforcement learning system can model and reason about persistent courses of action and perceive its environment in corresponding terms, and it develops and examines the validity of models of animal behavior related to this approach. The project's objectives are to develop the mathematical theory of the approach, to refine, extend, and conduct validation studies of related models of animal behavior, to examine the theory's relationship to control theory and artificial intelligence, and to demonstrate its effectiveness in a number of simulated learning tasks.Most current reinforcement learning (RL) research uses a framework in which an agent has to take a sequence of actions paced by a single, fixed time step: actions take one step to complete, and their immediate consequences become available after one step (modeled as a Markov decision process, or MDP). This makes it difficult to learn and plan at different time scales. Some RL research instead uses a generalization of this framework (semi-Markov decision processes, or SMDPs) in which actions take varying amounts of time to complete, and the existing theory specifies how to model the results of these actions and how to plan with them. However, this approach is limited because temporally extended actions are treated as indivisible and unknown units. For the greatest flexibility and best performance, it is necessary to look inside temporally extended actions to examine or modify how they are comprised of lower-level actions, which is not considered in existing approaches.This project, by contrast, will model extended courses of action as SMDP actions overlaid upon a base MDP. These courses of action, called options, can then be treated as if they were primitive actions; existing RL can be applied almost unchanged. This approach enables options to be analyzed at both the MDP and SMDP levels and introduces new issues at the interface between the levels. This approach is appealing because of its simplicity, similarity to previous approaches using primitive actions, and its solid mathematical foundation in MDP and SMDP theory. This is being developed further into a general approach to hierarchical and multi-time-scale planning and learning.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CRCNS: Collaborative Research: Neural Correlates of Hierarchical Reinforcement Learning
  • 批准号:
    1208051
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $4.33万
  • 财政年份:
    2012
  • 负责人:
    Andrew Barto
  • 依托单位:
NRI-Small: Collaborative Research: Multiple Task Learning from Unstructured Demonstrations
  • 批准号:
    1208497
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.92万
  • 财政年份:
    2012
  • 负责人:
    Andrew Barto
  • 依托单位:
SGER: Building Blocks for Creative Search
  • 批准号:
    0733581
  • 项目类别:
    Standard Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2007
  • 负责人:
    Andrew Barto
  • 依托单位:
Collaborative Research: Intrinsically Motivated Learning in Artificial Agents
  • 批准号:
    0432143
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2004
  • 负责人:
    Andrew Barto
  • 依托单位:
海外基金