Policy Transfer with Strategy Optimization

Policy Transfer with Strategy Optimization
复制标题

DOI:
--
复制
发表时间:
2018-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Wenhao Yu;C. Liu;Greg Turk
Wenhao Yu;C. Liu;Greg Turk
中科院分区:
其他
文献类型:
--
作者:
Wenhao Yu;C. Liu;Greg Turk

文献摘要

被引文献

相似文献

计算机仿真为训练机器人控制策略以完成运动等复杂任务提供了一种自动、安全的方法。然而,由于两种环境之间的差异,在模拟中训练的策略通常不会直接转移到实际硬件中。使用领域随机化的迁移学习是一种很有前途的方法,但它通常假设目标环境接近训练环境的分布,因此严重依赖于准确的系统识别。在本文中,我们提出了一种不同的方法,利用域随机化将控制策略转移到未知环境。关键的思想是,我们不是在模拟中学习单一的策略,而是同时学习一系列表现出不同行为的策略。在目标环境中进行测试时,我们直接根据任务性能在族中搜索最佳策略,而不需要识别动态参数。我们在训练和测试环境中对五个具有不同差异的模拟机器人控制问题进行了评估,并证明与训练鲁棒策略或自适应策略相比,我们的方法可以克服更大的建模误差。
Computer simulation provides an automatic and safe way for training robotic control policies to achieve complex tasks such as locomotion. However, a policy trained in simulation usually does not transfer directly to the real hardware due to the differences between the two environments. Transfer learning using domain randomization is a promising approach, but it usually assumes that the target environment is close to the distribution of the training environments, thus relying heavily on accurate system identification. In this paper, we present a different approach that leverages domain randomization for transferring control policies to unknown environments. The key idea that, instead of learning a single policy in the simulation, we simultaneously learn a family of policies that exhibit different behaviors. When tested in the target environment, we directly search for the best policy in the family based on the task performance, without the need to identify the dynamic parameters. We evaluate our method on five simulated robotic control problems with different discrepancies in the training and testing environment and demonstrate that our method can overcome larger modeling errors compared to training a robust policy or an adaptive policy.