Exploiting Structure in Reinforcement Learning Problems
Exploiting Structure in Reinforcement Learning Problems
批准号:
9711753
负责人:
Satinder Baveja
金额:
$22.97万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1997
资助国家:
美国
项目状态:
已结题
起止时间:
1997-12-01 至 1998-11-30
中文摘要
通过交互学习或强化学习的算法通常忽略环境中的所有结构,因此往往伸缩性较差。这项研究的目标是为结构化环境中的交互学习开发新的、高效的、理论上有充分基础的算法和架构。考虑了三种环境结构:状态和动作的因子结构、报酬函数的加性结构和状态和动作的层次结构。这种结构很常见,因为许多环境都是由多个交互较弱的组件组成的,这些组件通常是按层次结构组织的。该方法包括通过分别学习不同的组件来利用这种结构,然后以依赖于结构的方式对如此引入的近似进行补偿。这项研究的结果将阐明交互学习中常见的许多不同的有趣和有用的结构,并提供新的强化学习算法,使其能够解决比传统方法更大的结构化问题。可能的应用包括大规模的、动态的、面向通信、网络和调度的资源分配问题,以及来自分布式控制和人工智能的多智能体问题。
英文摘要
Algorithms for learning by interaction, or reinforcement learning, typically ignore all structure in the environment and consequently tend to scale poorly. The goal of this research is to develop novel, efficient, and theoretically well-founded algorithms and architectures for learning by interaction in structured environments. Three kinds of environmental structure are considered: factorial structure in states and actions, additive structure in payoff functions, and hierarchical structure in states and actions. Such structure is common because many environments are composed from multiple, weakly interacting, components that are often organized hierarchically. The approach consists of exploiting this structure by learning separately for the different components and then compensating in a structure dependent manner for the approximation so introduced. The results of this research will elucidate many different interesting and useful structures common in learning by interaction problems and provide new reinforcement learning algorithms that make it possible to solve significantly larger structured problems than possible with the traditional approach. Possible applications include large-scale, dynamic, resource allocation problems intelecommunications, networking, and scheduling, as well as multi-agent problems from distributed control and artificial intelligence.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Combining Reinforcement Learning and Deep Learning Methods to Address High-Dimensional Perception, Partial Observability and Delayed Reward
-
批准号:1526059
-
项目类别:Standard Grant
-
资助金额:$49.99万
-
财政年份:2015
-
负责人:Satinder Baveja
-
依托单位:
RI: Small: Reinforcement Learning with Predictive State Representations
-
批准号:1319365
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2013
-
负责人:Satinder Baveja
-
依托单位:
EAGER: On the Optimal Rewards Problem
-
批准号:1148668
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2011
-
负责人:Satinder Baveja
-
依托单位:
SHB: Medium: Collaborative Research: Novel Computational Techniques for Cardiovascular Risk Stratification
-
批准号:1064948
-
项目类别:Standard Grant
-
资助金额:$56.24万
-
财政年份:2011
-
负责人:Satinder Baveja
-
依托单位:
RI: Medium: Building Flexible, Robust, and Autonomous Agents
-
批准号:0905146
-
项目类别:Standard Grant
-
资助金额:$120.0万
-
财政年份:2009
-
负责人:Satinder Baveja
-
依托单位:
Flexible State Representations in Reinforcement Learning
-
批准号:0413004
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Satinder Baveja
-
依托单位:
Collaborative Research: Intrinsically Motivated Learning in Artificial Agents
-
批准号:0432027
-
项目类别:Continuing Grant
-
资助金额:$15.0万
-
财政年份:2004
-
负责人:Satinder Baveja
-
依托单位:
海外基金