Towards General and Autonomous Learning of Core Skills: A Case Study in Locomotion

Towards General and Autonomous Learning of Core Skills: A Case Study in Locomotion
复制标题

DOI:
--
复制
发表时间:
2020-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Roland Hafner;Tim Hertweck;Philipp Kloppner;Michael Bloesch;Michael Neunert;Markus Wulfmeier;S. Tunyasuvunakool;N. Heess;Martin A. Riedmiller
Roland Hafner;Tim Hertweck;Philipp Kloppner;Michael Bloesch;Michael Neunert;Markus Wulfmeier;S. Tunyasuvunakool;N. Heess;Martin A. Riedmiller
中科院分区:
其他
文献类型:
--
作者:
Roland Hafner;Tim Hertweck;Philipp Kloppner;Michael Bloesch;Michael Neunert;Markus Wulfmeier;S. Tunyasuvunakool;N. Heess;Martin A. Riedmiller

文献摘要

被引文献

相似文献

现代强化学习(RL)算法有望直接从原始感官输入解决困难的运动控制问题。它们的吸引力部分是由于它们可以代表一类一般的方法,这些方法允许用合理的奖励和最小的先验知识来学习解决方案,即使在对人类专家来说困难或昂贵的情况下也是如此。然而,为了让强化学习真正实现这一承诺,我们需要算法和学习设置,这些算法和学习设置可以通过最小的问题特定调整或工程来解决广泛的问题。在本文中,我们在运动领域研究了这种一般性思想。我们开发了一个学习框架,可以学习广泛的有腿机器人的复杂运动行为,如两足动物、三足动物、四足动物和六足动物,包括轮式变体。我们的学习框架依赖于一个数据高效、非策略的多任务强化学习算法和一组在机器人之间语义相同的奖励函数。为了强调该方法的一般适用性,我们在实验中保持超参数设置和奖励定义不变,并完全依赖于车载传感。对于九种不同类型的机器人,包括现实世界的四足机器人,我们证明了相同的算法可以快速学习各种可重复使用的运动技能,而无需任何平台特定的调整或额外的学习设置仪器。
Modern Reinforcement Learning (RL) algorithms promise to solve difficult motor control problems directly from raw sensory inputs. Their attraction is due in part to the fact that they can represent a general class of methods that allow to learn a solution with a reasonably set reward and minimal prior knowledge, even in situations where it is difficult or expensive for a human expert. For RL to truly make good on this promise, however, we need algorithms and learning setups that can work across a broad range of problems with minimal problem specific adjustments or engineering. In this paper, we study this idea of generality in the locomotion domain. We develop a learning framework that can learn sophisticated locomotion behavior for a wide spectrum of legged robots, such as bipeds, tripeds, quadrupeds and hexapods, including wheeled variants. Our learning framework relies on a data-efficient, off-policy multi-task RL algorithm and a small set of reward functions that are semantically identical across robots. To underline the general applicability of the method, we keep the hyper-parameter settings and reward definitions constant across experiments and rely exclusively on on-board sensing. For nine different types of robots, including a real-world quadruped robot, we demonstrate that the same algorithm can rapidly learn diverse and reusable locomotion skills without any platform specific adjustments or additional instrumentation of the learning setup.