课题基金 / 基金详情

KDI: Temporal Abstraction in Reinforcement Learning

KDI: Temporal Abstraction in Reinforcement Learning
KDI:强化学习中的时间抽象
批准号:
9980062
负责人:
Andrew Barto
金额:
$56.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
1999
资助国家:
美国
项目状态:
已结题
起止时间:
1999-09-15 至 2003-08-31

项目摘要

项目成果

Andrew Barto的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
This project investigates a new approach to learning, planning, and representing knowledge at multiple levels of temporal abstraction. It develops methods by which an artificial reinforcement learning system can model and reason about persistent courses of action and perceive its environment in corresponding terms, and it develops and examines the validity of models of animal behavior related to this approach. The project's objectives are to develop the mathematical theory of the approach, to refine, extend, and conduct validation studies of related models of animal behavior, to examine the theory's relationship to control theory and artificial intelligence, and to demonstrate its effectiveness in a number of simulated learning tasks.Most current reinforcement learning (RL) research uses a framework in which an agent has to take a sequence of actions paced by a single, fixed time step: actions take one step to complete, and their immediate consequences become available after one step (modeled as a Markov decision process, or MDP). This makes it difficult to learn and plan at different time scales. Some RL research instead uses a generalization of this framework (semi-Markov decision processes, or SMDPs) in which actions take varying amounts of time to complete, and the existing theory specifies how to model the results of these actions and how to plan with them. However, this approach is limited because temporally extended actions are treated as indivisible and unknown units. For the greatest flexibility and best performance, it is necessary to look inside temporally extended actions to examine or modify how they are comprised of lower-level actions, which is not considered in existing approaches.This project, by contrast, will model extended courses of action as SMDP actions overlaid upon a base MDP. These courses of action, called options, can then be treated as if they were primitive actions; existing RL can be applied almost unchanged. This approach enables options to be analyzed at both the MDP and SMDP levels and introduces new issues at the interface between the levels. This approach is appealing because of its simplicity, similarity to previous approaches using primitive actions, and its solid mathematical foundation in MDP and SMDP theory. This is being developed further into a general approach to hierarchical and multi-time-scale planning and learning.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CRCNS: Collaborative Research: Neural Correlates of Hierarchical Reinforcement Learning
  • 批准号:
    1208051
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $4.33万
  • 财政年份:
    2012
  • 负责人:
    Andrew Barto
  • 依托单位:
NRI-Small: Collaborative Research: Multiple Task Learning from Unstructured Demonstrations
  • 批准号:
    1208497
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.92万
  • 财政年份:
    2012
  • 负责人:
    Andrew Barto
  • 依托单位:
SGER: Building Blocks for Creative Search
  • 批准号:
    0733581
  • 项目类别:
    Standard Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2007
  • 负责人:
    Andrew Barto
  • 依托单位:
Collaborative Research: Intrinsically Motivated Learning in Artificial Agents
  • 批准号:
    0432143
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2004
  • 负责人:
    Andrew Barto
  • 依托单位:
海外基金