Empowered skills

Empowered skills
复制标题

DOI:
10.1109/icra.2017.7989760
复制
发表时间:
2017-05
期刊:
2017 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Alexander Gabriel;R. Akrour;Jan Peters;G. Neumann
Alexander Gabriel;R. Akrour;Jan Peters;G. Neumann
中科院分区:
其他
文献类型:
--
作者:
Alexander Gabriel;R. Akrour;Jan Peters;G. Neumann

文献摘要

相似文献

机器人强化学习(RL)算法返回一个最大化全局累积奖励信号的策略,但通常不会创建不同的行为。因此,策略通常只捕获任务的单个解决方案。然而,许多运动任务有各种各样的解决方案,关于这些解决方案的知识可以有几个优点。例如,在机器人乒乓球这样的对抗性环境中,缺乏多样性使得行为可预测,因此对手很容易反击。在交互式环境中,例如从人类反馈中学习,对多样性的强调使人类有更多的机会指导机器人,并避免后者陷入任务的局部最优。为了增加学到的行为的多样性,我们利用以前的工作内在动机和授权。我们得到一个新的内在动机信号,丰富的描述与结果空间的任务,代表感兴趣的方面的运动流。例如,在乒乓球中,结果空间可以由回球位置和回球速度给出。现在,内在动机是未来成果的多样性,这一概念也被称为增强权能。我们推导出一个新的政策搜索算法,最大限度地权衡外在的奖励和内在的动机标准。在平面到达任务和模拟机器人乒乓球上的实验表明,我们的算法可以学习任务感兴趣区域内的各种行为。
Robot Reinforcement Learning (RL) algorithms return a policy that maximizes a global cumulative reward signal but typically do not create diverse behaviors. Hence, the policy will typically only capture a single solution of a task. However, many motor tasks have a large variety of solutions and the knowledge about these solutions can have several advantages. For example, in an adversarial setting such as robot table tennis, the lack of diversity renders the behavior predictable and hence easy to counter for the opponent. In an interactive setting such as learning from human feedback, an emphasis on diversity gives the human more opportunity for guiding the robot and to avoid the latter to be stuck in local optima of the task. In order to increase diversity of the learned behaviors, we leverage prior work on intrinsic motivation and empowerment. We derive a new intrinsic motivation signal by enriching the description of a task with an outcome space, representing interesting aspects of a sensorimotor stream. For example, in table tennis, the outcome space could be given by the return position and return ball speed. The intrinsic motivation is now given by the diversity of future outcomes, a concept also known as empowerment. We derive a new policy search algorithm that maximizes a trade-off between the extrinsic reward and this intrinsic motivation criterion. Experiments on a planar reaching task and simulated robot table tennis demonstrate that our algorithm can learn a diverse set of behaviors within the area of interest of the tasks.