HARL-A: Hardware Agnostic Reinforcement Learning Through Adversarial Selection

HARL-A: Hardware Agnostic Reinforcement Learning Through Adversarial Selection
复制标题

DOI:
10.1109/iros51168.2021.9636167
复制
发表时间:
2021-09
期刊:
2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Lucy Jackson;S. Eckersley;Pete Senior;Simon Hadfield
Lucy Jackson;S. Eckersley;Pete Senior;Simon Hadfield
中科院分区:
其他
文献类型:
--
作者:
Lucy Jackson;S. Eckersley;Pete Senior;Simon Hadfield

文献摘要

相似文献

强化学习(RL)的使用导致了机器人领域的巨大进步。然而,数据稀缺,脆弱的收敛以及模拟与真实的世界环境之间的差距,意味着大多数常见的RL方法都会受到过度拟合的影响,并且无法推广到不可见的环境。硬件不可知策略将通过允许单个网络在各种测试域中运行来缓解这一问题,其中动态会因机器人形态或内部参数的变化而变化。我们利用的想法,学习适应一个已知的和成功的控制策略是更容易和更灵活的,比联合学习众多的控制策略,为不同的morphology.This本文提出的思想,硬件不可知的强化学习使用对抗性选择(HARL-A)。在这种方法中,使用一种新的对抗性损失函数对训练样本进行采样。这是为了根据它们的学习潜力来自我调节形态。简单地将我们的基于学习潜力的损失函数应用于当前最先进的技术已经提供了约30%的性能改进。与此同时,使用HARL-A的完整实现的实验报告,与标准RL基线相比平均增加了70%,与当前最先进的技术相比平均增加了55%。
The use of reinforcement learning (RL) has led to huge advancements in the field of robotics. However data scarcity, brittle convergence and the gap between simulation & real world environments, mean that most common RL approaches are subject to over fitting and fail to generalise to unseen environments. Hardware agnostic policies would mitigate this by allowing a single network to operate in a variety of test domains, where dynamics vary due to changes in robotic morphologies or internal parameters. We utilise the idea that learning to adapt a known and successful control policy is easier and more flexible than jointly learning numerous control policies for different morphologies.This paper presents the idea of Hardware Agnostic Reinforcement Learning using Adversarial selection (HARL-A). In this approach training examples are sampled using a novel adversarial loss function. This is designed to self regulate morphologies based on their learning potential. Simply applying our learning potential based loss function to current state-of-the-art already provides ~ 30% improvement in performance. Meanwhile experiments using the full implementation of HARL-A report an average increase of 70% to a standard RL baseline and 55% compared with current state-of-the-art.