Scaling Reinforcement Learning by Adaptive Task Selection and Linear Solution Merging
Scaling Reinforcement Learning by Adaptive Task Selection and Linear Solution Merging
批准号:
9501852
负责人:
Sridhar Mahadevan
金额:
$17.41万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1995
资助国家:
美国
项目状态:
已结题
起止时间:
1995-06-01 至 1998-05-31
中文摘要
所提出的研究的目的是研究自治代理如何能够适应动态部分已知的任务环境。 这种代理的潜在应用范围从自动交付家务的硬件机器人到从互联网检索信息的软件程序。 本研究将集中在自适应控制 这就是所谓的强化学习 在这种方法中,智能体通过试验和错误来获得任务技能 通过 选择 行动 的 最大 一 奖励函数。强化学习有一些问题。 它的收敛速度非常慢,特别是在奖励很少出现的大状态空间问题中。此外,学习到的技能在相关任务之间的转移也很差。 本研究将研究使用一种新的模块化任务架构来克服强化学习的这些局限性。该架构将复合多目标任务分解为实现每个目标的原始子任务。 它利用训练时间更多 通过基于任务的难度和重要性在学习不同任务之间动态切换来有效地学习。 它通过使用加权线性和函数将学习到的解决方案重用到原始子任务来增加任务之间的传输。 它通过使用强化学习方法来更有效地解决重复性任务, 奖励 一个详细的实验研究的建议 架构 将 进行,使用各种模拟和真实的机器人试验台。
英文摘要
The aim of the proposed research is to study how autonomous agents can adapt to dynamic partially known task environments. Potential applications of such agents range from hardware robots that automate delivery chores to software programs that retrieve information from the Internet. This research will focus on an adaptive control paradigm called reinforcement learning. In this approach, agents acquire task skills through trial and error by selecting actions that maximize a reward function.Reinforcement learning has some problems. It converges extremely slowly, especially in large state space problems where rewards occur infrequently. Also, the learned skills transfer poorly across related tasks. This research will investigate using a novel modular task architecture to overcome these limitations of reinforcement learning. The proposed architecture decomposes composite multiple goal tasks into primitive subtasks that achieve each individual goal. It utilizes training time more efficiently by dynamically switching between learning different tasks based on their difficulty and importance. It increases transfer across tasks by reusing solutions learned to primitive subtasks using a weighted linear sum function. It solves recurrent tasks more effectively by using a reinforcement learning method that optimizes average reward. A detailed experimental study of the proposed architecture will be undertaken, using a variety of simulated and real robot testbeds.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Transfer Learning for Chemical Analyses from Laser-Induced Spectroscopy
-
批准号:1307179
-
项目类别:Standard Grant
-
资助金额:$15.89万
-
财政年份:2013
-
负责人:Sridhar Mahadevan
-
依托单位:
RI: Small: Reinforcement Learning by Mirror Descent
-
批准号:1216467
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2012
-
负责人:Sridhar Mahadevan
-
依托单位:
NeTS Small: Analysis and Design of Best-Effort Content-Caching Networks
-
批准号:1117764
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2011
-
负责人:Sridhar Mahadevan
-
依托单位:
Manifold Alignment of High-Dimensional Data Sets
-
批准号:1025120
-
项目类别:Standard Grant
-
资助金额:$49.99万
-
财政年份:2010
-
负责人:Sridhar Mahadevan
-
依托单位:
RI-Medium: Collaborative Research: Learning Multiscale Representations using Harmonic Analysis on Graphs
-
批准号:0803288
-
项目类别:Standard Grant
-
资助金额:$34.52万
-
财政年份:2008
-
负责人:Sridhar Mahadevan
-
依托单位:
Proto-Value Functions: A Unified Framework for Learning Task-Specific Behaviors and Task-Independent Representations
-
批准号:0534999
-
项目类别:Continuing Grant
-
资助金额:$44.36万
-
财政年份:2006
-
负责人:Sridhar Mahadevan
-
依托单位:
Scaling Reinforcement Learning by Adaptive Task Selection and Linear Solution Merging
-
批准号:9896122
-
项目类别:Continuing Grant
-
资助金额:$7.85万
-
财政年份:1997
-
负责人:Sridhar Mahadevan
-
依托单位:
Support for a Workshop on Reinforcement Learning
-
批准号:9529108
-
项目类别:Standard Grant
-
资助金额:$2.96万
-
财政年份:1995
-
负责人:Sridhar Mahadevan
-
依托单位:
国内基金
海外基金
海桑属杂种区强化(Reinforcement)的检验与遗传基础研究
-
批准号:30800060
-
项目类别:青年科学基金项目
-
资助金额:23.0万元
-
批准年份:2008
-
负责人:周仁超
-
依托单位: