Scalable Autonomous Reinforcement Learning - From scratch to less and less structure
Scalable Autonomous Reinforcement Learning - From scratch to less and less structure
批准号:
260194412
负责人:
Professor Dr. Joschka Bödecker, since 4/2015
金额:
$0.0万
依托单位国家:
德国
项目类别:
Priority Programmes
财政年份:
2014
资助国家:
德国
项目状态:
已结题
起止时间:
2013-12-31 至 2020-12-31
中文摘要
在过去的十年中,强化学习框架(RL)已经发展成为一种有前途的工具,用于学习机器人中各种不同的任务。在此期间,在将强化学习扩展到高维系统和解决日益复杂的任务方面取得了很多进展。不幸的是,这种可扩展性是通过使用专家知识在几个维度上预先构建学习问题来实现的。因此,机器人强化学习中最先进的方法通常依赖于手工制作的状态表示、预结构的参数化策略、形状良好的奖励函数和人类专家的演示,以帮助扩展学习算法。在这个提议中,我们希望通过一个具有挑战性的机器人任务(即绳球)的“经典”强化学习设置来推进该领域。通过强化学习方法解决这个任务将是一个有价值的贡献。从这里开始,我们将开始确定学习任务设计仍然需要工程经验的组件。在本建议的过程中,我们将展示如何在开发高度可扩展的方法的同时,推动这些组件实现更多的自主权。为此,我们将开发系统的方法,通过超越传统的方法来增加学习系统的自主性:(1)提出自动学习强化学习状态表示的方法;(2)开发能够表示真正自主行为所必需的大量控制策略的通用策略类;(3)自主发现信息性奖励函数。这些方面的进展将把学习算法提升到更高的自治水平。这些进步将以完善的政策搜索理论框架为基础,并通过改进最先进的强化学习算法来实现。最终产生的系统应该学会如何从简单、通用的原则中将原始感官输入映射到原始控制信号,自动发现其环境中的结构,并在没有专家知识的情况下解决困难的控制任务。如果成功,该项目中开发的完整方法及其子部分将有助于建立一个新的,更强大的一代强化学习算法,能够自主解决复杂的机器人控制问题。
英文摘要
Over the course of the last decade, the framework of reinforcement learning (RL) has developed into a promising tool for learning a large variety of different tasks in robotics. During this timeframe, a lot of progress has been made towards scaling reinforcement learning to high-dimensional systems and solving tasks of increasing complexity. Unfortunately, this scalability has been achieved by using expert knowledge to pre-structure the learning problem in several dimensions. As a consequence, the state-of-the-art methods in robot reinforcement learning generally depend on hand-crafted state representations, pre-structured parametrized policies, well-shaped reward functions and demonstrations by a human expert to aid scaling of the learning algorithm.In this proposal, we want to advance the field by starting with a 'classical' reinforcement learning setting for a challenging robotic task (i.e., tetherball). Solving this task by RL methods will be already a valuable contribution. From there on, we will start to identify the components for which the learning task design still needs engineering experience. In the course of this proposal, we show how we aim to drive each of these components towards more autonomy while developing highly scalable approaches.To this end, we will develop systematic methods to increase the autonomy of the learning system by going beyond traditional approaches: (1) proposing methods for learning state representations for reinforcement learning automatically; (2) developing generic policy classes capable of representing the large variety of control policies that are necessary for truly autonomous behavior; (3) discovering informative reward functions autonomously. Progress in each of these aspects will lift the learning algorithm to a higher level of autonomy. The advances will be grounded in the well established theoretical framework of policy search and enabled through improvements to state-of-the-art reinforcement learning algorithms. Ultimately the resulting system should learn how to map raw sensory inputs to raw control signals from simple, generic principles, discovering structure within its environment automatically and solving difficult control tasks without expert knowledge. If successful, both the complete methodology developed within this project as well as sub-parts of it will help to establish a new, substantially more powerful generation of reinforcement learning algorithms that are capable of solving complicated robot control problems autonomously.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1016/j.neucom.2016.11.094
发表时间:
2017-11
期刊:
Neurocomputing
影响因子:
6
作者:
[Simone Parisi;Matteo Pirotta;Jan Peters]
通讯作者:
Simone Parisi;Matteo Pirotta;Jan Peters
DOI:
10.1109/iros.2015.7354296
发表时间:
2015-12
期刊:
2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
作者:
[Simone Parisi;Hany Abdulsamad;A. Paraschos;Christian Daniel;Jan Peters]
通讯作者:
Simone Parisi;Hany Abdulsamad;A. Paraschos;Christian Daniel;Jan Peters
DOI:
10.1109/iros.2017.8206334
发表时间:
2017-09
期刊:
2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
作者:
[Simone Parisi;Simon Ramstedt;Jan Peters]
通讯作者:
Simone Parisi;Simon Ramstedt;Jan Peters
Local-utopia policy selection for multi-objective reinforcement learning
多目标强化学习的本地乌托邦策略选择
DOI:
10.1109/ssci.2016.7849369
发表时间:
2016
期刊:
2016 IEEE Symposium Series on Computational Intelligence (SSCI)
影响因子:
--
作者:
[Simone Parisi, Alexander Blank, Tobias Viernickel, Jan Peters]
通讯作者:
Jan Peters
海外基金