课题基金 / 基金详情

Average Reward Reinforcement Learning: Scaling up

Average Reward Reinforcement Learning: Scaling up
平均奖励强化学习:扩大规模
批准号:
0098050
负责人:
Prasad Tadepalli
金额:
$36.87万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2001
资助国家:
美国
项目状态:
已结题
起止时间:
2001-05-01 至 2005-04-30

项目摘要

项目成果

Prasad Tadepalli的其他基金

相似基金

相关文献

中文摘要
翻译
这是三年持续奖的第一年资助。想象一下,未来的工厂充满了智能自主机器人和从事生产的机器。机器不仅可以感知和行动,而且还可以在没有明确编程的情况下优化自己的行为。以及能够学习与其他机器人协调行动以满足整体优化标准的机器人。制造这样的机器和机器人的能力从根本上重组了工厂,因此人的角色被简化为指定优化标准并向机器提供反馈,将低级别的控制和优化问题留给机器本身。这个项目旨在设计和研究制造这种机器和机器人所需的算法和计算工具。长期的科学目标是更好地理解自适应自主多智能体系统设计中涉及的权衡,特别是行为的最佳性、计算和通信效率、通用性和学习速度之间的权衡。通过奖惩或强化学习来优化程序的性能,似乎是为复杂的现实世界领域构建这种自适应多代理系统的最有前途的方法。制造业中的许多现实问题,如生产调度和库存控制,都可以看作是“平均报酬强化学习”(ARL)问题,其中的优化准则是最大化每单位时间获得的平均报酬。该项目的目标是开发可扩展的算法、程序和技术,以解决大型ARL问题,并将制造业作为主要应用领域。PI将推动这项技术的前沿,使其可以应用于拥有数百台机器和作业类型的工厂,并具有部分可观察性和对多个代理的可伸缩性等现实假设。这个项目的成功完成将带来新的可扩展的算法和程序来解决大型ARL问题,这很可能会产生重大的经济影响。
英文摘要
This is the first year funding of a three year continuing award. Imagine a factory of the future that teems with intelligent autonomous robots and machines engaged in production. Machines that can not only sense and act, but can also optimize their own behavior without being explicitly programmed. And robots that can learn to coordinate their actions with other robots in order to satisfy an overall optimization criterion. The ability to build such machines and robots radically reorganizes the factories, so that people's role is reduced to specifying an optimization criterion and giving feedback to the machines, leaving the low level control and optimization issues to the machines themselves. This project seeks to design and study the algorithmic and computational tools necessary to build such machines and robots. The long-term scientific goal is to gain a better understanding of the tradeoffs involved in the design of adaptive autonomous multi-agent systems; in particular, the tradeoffs between the optimality of the behavior, computational and communication efficiencies, generality, and speed of learning. Optimizing the performance of programs via rewards and punishments, or reinforcement learning, appears to be the most promising approach to building such adaptive multi-agent systems for complex real-world domains. Many real-world problems in manufacturing, such as production scheduling and inventory control, are best seen as "average-reward reinforcement learning" (ARL) problems, where the optimization criterion is to maximize the average reward received per unit of time. The goal of this project is to develop scaleable algorithms, programs and techniques for solving large ARL problems, with manufacturing as the primary application domain. The PI will push the frontiers of this technology to the point where it can be applied to factories with hundreds of machines and job types, with realistic assumptions such as partial observability and scalability to multiple agents. Successful completion of this project will lead to new scaleable algorithms and programs for solving large ARL problems, which could well have significant economic impact.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Integrating Learning and Search for Structured Prediction
  • 批准号:
    1219258
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2012
  • 负责人:
    Prasad Tadepalli
  • 依托单位:
RI: Medium: Collaborative Research: Optimizing Policies for Service Organizations in Complex Structured Domains
  • 批准号:
    0964705
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $58.17万
  • 财政年份:
    2010
  • 负责人:
    Prasad Tadepalli
  • 依托单位:
Relational Reinforcement Learning
  • 批准号:
    0329278
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $41.16万
  • 财政年份:
    2003
  • 负责人:
    Prasad Tadepalli
  • 依托单位:
Average Reward Reinforcement Learning
  • 批准号:
    9520243
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $22.49万
  • 财政年份:
    1995
  • 负责人:
    Prasad Tadepalli
  • 依托单位:
海外基金