Robust Learning from Observation with Model Misspecification

Robust Learning from Observation with Model Misspecification
复制标题

DOI:
10.5555/3535850.3535999
复制
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Luca Viano;Yu-ting Huang;Parameswaran Kamalaruban;Craig Innes;S. Ramamoorthy;Adrian Weller
Luca Viano;Yu-ting Huang;Parameswaran Kamalaruban;Craig Innes;S. Ramamoorthy;Adrian Weller
中科院分区:
其他
文献类型:
--
作者:
Luca Viano;Yu-ting Huang;Parameswaran Kamalaruban;Craig Innes;S. Ramamoorthy;Adrian Weller

文献摘要

被引文献

相似文献

当难以确定奖励函数时,模仿学习(IL)是机器人系统中训练策略的一种流行范例。然而,尽管IL算法取得了成功,但它们提出了一些不切实际的要求,即专家演示必须来自要学习新的模仿者策略的同一领域。我们考虑一个实际的设置,其中(i)向学习者提供来自真实(部署)环境的纯状态专家演示,(ii)模仿学习者策略在模拟(训练)环境中进行训练,其过渡动态与真实环境略有不同,以及(iii)学习者在训练阶段没有任何访问真实环境的机会,超出了给出的演示批次。目前的大多数IL方法,如生成对抗模仿学习及其状态变量,都无法模仿上述设置下的最优专家行为。通过利用鲁棒强化学习(RL)文献的见解并基于最近的对抗性模仿方法,我们提出了一种鲁棒强化学习算法来学习可以有效地转移到真实环境而无需微调的策略。此外,我们在连续控制基准上的经验证明,我们的方法在真实环境中的零射击传递性能和不同测试条件下的鲁棒性能方面优于最先进的仅状态IL方法。
Imitation learning (IL) is a popular paradigm for training policies in robotic systems when specifying the reward function is difficult. However, despite the success of IL algorithms, they impose the somewhat unrealistic requirement that the expert demonstrations must come from the same domain in which a new imitator policy is to be learned. We consider a practical setting, where (i) state-only expert demonstrations from the real (deployment) environment are given to the learner, (ii) the imitation learner policy is trained in a simulation (training) environment whose transition dynamics is slightly different from the real environment, and (iii) the learner does not have any access to the real environment during the training phase beyond the batch of demonstrations given. Most of the current IL methods, such as generative adversarial imitation learning and its state-only variants, fail to imitate the optimal expert behavior under the above setting. By leveraging insights from the Robust reinforcement learning (RL) literature and building on recent adversarial imitation approaches, we propose a robust IL algorithm to learn policies that can effectively transfer to the real environment without fine-tuning. Furthermore, we empirically demonstrate on continuous-control benchmarks that our method outperforms the state-of-the-art state-only IL method in terms of the zero-shot transfer performance in the real environment and robust performance under different testing conditions.