Distributionally Robust Behavioral Cloning for Robust Imitation Learning

Distributionally Robust Behavioral Cloning for Robust Imitation Learning
复制标题

DOI:
10.1109/cdc49753.2023.10383976
复制
发表时间:
2023-12
期刊:
2023 62nd IEEE Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
Kishan Panaganti;Zaiyan Xu;D. Kalathil;Mohammad Ghavamzadeh
Kishan Panaganti;Zaiyan Xu;D. Kalathil;Mohammad Ghavamzadeh
中科院分区:
其他
文献类型:
--
作者:
Kishan Panaganti;Zaiyan Xu;D. Kalathil;Mohammad Ghavamzadeh

文献摘要

相似文献

鲁棒强化学习 (RL) 旨在学习一种能够承受模型参数不确定性的策略,这种不确定性通常在实际 RL 应用中由于模拟器中的建模错误、现实系统动力学的变化和对抗性干扰而出现。本文介绍了马尔可夫决策过程 (MDP) 框架中的鲁棒模仿学习 (IL) 问题,其中代理学习模仿专家演示器,该演示器可以承受模型参数的不确定性,而无需额外的在线环境交互。代理仅获得来自专家的关于单个(名义)动态的状态-动作对数据集,而没有关于环境中真实奖励的任何信息。行为克隆 (BC) 是一种监督学习方法,是解决普通 IL 问题的强大算法。我们提出了一种鲁棒 IL 问题的算法,该算法利用 BC 的分布鲁棒优化 (DRO)。我们将该算法称为 DR-BC,并在理论和实践中展示了其针对参数不确定性的鲁棒性能。我们还展示了我们的方法在几个 MuJoCo 连续控制任务上解决模型扰动的经验性能。
Robust reinforcement learning (RL) aims to learn a policy that can withstand uncertainties in model parameters, which often arise in practical RL applications due to modeling errors in simulators, variations in real-world system dynamics, and adversarial disturbances. This paper introduces the robust imitation learning (IL) problem in a Markov decision process (MDP) framework where an agent learns to mimic an expert demonstrator that can withstand uncertainties in model parameters without additional online environment interactions. The agent is only provided with a dataset of state-action pairs from the expert on a single (nominal) dynamics, without any information about the true rewards from the environment. Behavioral cloning (BC), a supervised learning method, is a powerful algorithm to address the vanilla IL problem. We propose an algorithm for the robust IL problem that utilizes distributionally robust optimization (DRO) with BC. We call the algorithm DR-BC and show its robust performance against parameter uncertainties both in theory and in practice. We also demonstrate the empirical performance of our approach to addressing model perturbations on several MuJoCo continuous control tasks.