课题基金 / 基金详情

Exploiting Structure in Reinforcement Learning Problems

Exploiting Structure in Reinforcement Learning Problems
利用强化学习问题中的结构
批准号:
9711753
负责人:
Satinder Baveja
金额:
$22.97万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1997
资助国家:
美国
项目状态:
已结题
起止时间:
1997-12-01 至 1998-11-30

项目摘要

项目成果

Satinder Baveja的其他基金

相似基金

相关文献

中文摘要
翻译
通过交互学习或强化学习的算法通常忽略环境中的所有结构,因此往往伸缩性较差。这项研究的目标是为结构化环境中的交互学习开发新的、高效的、理论上有充分基础的算法和架构。考虑了三种环境结构:状态和动作的因子结构、报酬函数的加性结构和状态和动作的层次结构。这种结构很常见,因为许多环境都是由多个交互较弱的组件组成的,这些组件通常是按层次结构组织的。该方法包括通过分别学习不同的组件来利用这种结构,然后以依赖于结构的方式对如此引入的近似进行补偿。这项研究的结果将阐明交互学习中常见的许多不同的有趣和有用的结构,并提供新的强化学习算法,使其能够解决比传统方法更大的结构化问题。可能的应用包括大规模的、动态的、面向通信、网络和调度的资源分配问题,以及来自分布式控制和人工智能的多智能体问题。
英文摘要
Algorithms for learning by interaction, or reinforcement learning, typically ignore all structure in the environment and consequently tend to scale poorly. The goal of this research is to develop novel, efficient, and theoretically well-founded algorithms and architectures for learning by interaction in structured environments. Three kinds of environmental structure are considered: factorial structure in states and actions, additive structure in payoff functions, and hierarchical structure in states and actions. Such structure is common because many environments are composed from multiple, weakly interacting, components that are often organized hierarchically. The approach consists of exploiting this structure by learning separately for the different components and then compensating in a structure dependent manner for the approximation so introduced. The results of this research will elucidate many different interesting and useful structures common in learning by interaction problems and provide new reinforcement learning algorithms that make it possible to solve significantly larger structured problems than possible with the traditional approach. Possible applications include large-scale, dynamic, resource allocation problems intelecommunications, networking, and scheduling, as well as multi-agent problems from distributed control and artificial intelligence.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Combining Reinforcement Learning and Deep Learning Methods to Address High-Dimensional Perception, Partial Observability and Delayed Reward
RI: Small: Reinforcement Learning with Predictive State Representations
EAGER: On the Optimal Rewards Problem
SHB: Medium: Collaborative Research: Novel Computational Techniques for Cardiovascular Risk Stratification
海外基金