Distributionally Robust Policy Learning via Adversarial Environment Generation

Distributionally Robust Policy Learning via Adversarial Environment Generation
复制标题

DOI:
10.1109/lra.2021.3139949
复制
发表时间:
2021-07
影响因子:
5.2
通讯作者:
Allen Z. Ren;Anirudha Majumdar
Allen Z. Ren;Anirudha Majumdar
中科院分区:
计算机科学2区
文献类型:
--
作者:
Allen Z. Ren;Anirudha Majumdar

文献摘要

相似文献

我们的目标是训练控制策略,使其能够很好地推广到不可见的环境。受分布式鲁棒优化(DRO)框架的启发,我们提出了DRAGEN -通过对抗性生成增强的分布式鲁棒策略学习-通过生成对抗性环境来迭代地提高策略对现实分布变化的鲁棒性。其关键思想是学习环境的生成模型,其潜变量捕获环境中的成本预测和现实变化。我们执行DRO相对于Wasserstein球周围的环境的经验分布通过生成现实的对抗性环境,通过梯度上升的潜在空间。我们在模拟中展示了强大的分布外(OoD)泛化,用于(i)用板载视觉摆动钟摆和(ii)抓住逼真的3D物体。硬件上的抓取实验表明,与域随机化相比,sim 2 real具有更好的性能。
Our goal is to train control policies that generalize well to unseen environments. Inspired by the Distributionally Robust Optimization (DRO) framework, we propose DRAGEN — Distributionally Robust policy learning via Adversarial Generation of ENvironments — for iteratively improving robustness of policies to realistic distribution shifts by generating adversarial environments. The key idea is to learn a generative model for environments whose latent variables capture cost-predictive and realistic variations in environments. We perform DRO with respect to a Wasserstein ball around the empirical distribution of environments by generating realistic adversarial environments via gradient ascent on the latent space. We demonstrate strong Out-of-Distribution (OoD) generalization in simulation for (i) swinging up a pendulum with onboard vision and (ii) grasping realistic 3D objects. Grasping experiments on hardware demonstrate better sim2real performance compared to domain randomization.