Self-Adaptive Imitation Learning: Learning Tasks with Delayed Rewards from Sub-optimal Demonstrations

Self-Adaptive Imitation Learning: Learning Tasks with Delayed Rewards from Sub-optimal Demonstrations
复制标题

DOI:
10.1609/aaai.v36i8.20914
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Zhuangdi Zhu;Kaixiang Lin;Bo Dai;Jiayu Zhou
Zhuangdi Zhu;Kaixiang Lin;Bo Dai;Jiayu Zhou
中科院分区:
其他
文献类型:
--
作者:
Zhuangdi Zhu;Kaixiang Lin;Bo Dai;Jiayu Zhou

文献摘要

被引文献

相似文献

强化学习(RL)在解决序贯决策问题中显示了其优越性。然而,对即时奖赏反馈的严重依赖阻碍了RL的广泛应用。另一方面,模仿学习(IL)通过利用外部示范在不依赖环境监督的情况下解决RL。然而,在实践中,收集足够的专家演示可能昂贵得令人望而却步,但演示的质量通常会限制学习政策的执行。为了解决一个实际情况,在这项工作中,我们提出了自适应模仿学习(SAIL),它由一个次优教师提供的一些演示提供,可以在具有极大延迟回报的RL任务中表现良好,其中唯一的回报反馈是按轨迹排序。SAIL通过互动地利用演示来赶上老师,并探索环境以产生超过老师的演示,从而将IL和RL的优势连接起来。广泛的实证结果表明,SAIL不仅显著提高了样本效率,而且与最先进的连续控制任务相比,它还导致了不同连续控制任务的更高的渐近性能。
Reinforcement learning (RL) has demonstrated its superiority in solving sequential decision-making problems. However, heavy dependence on immediate reward feedback impedes the wide application of RL. On the other hand, imitation learning (IL) tackles RL without relying on environmental supervision by leveraging external demonstrations. In practice, however, collecting sufficient expert demonstrations can be prohibitively expensive, yet the quality of demonstrations typically limits the performance of the learning policy. To address a practical scenario, in this work, we propose Self-Adaptive Imitation Learning (SAIL), which, provided with a few demonstrations from a sub-optimal teacher, can perform well in RL tasks with extremely delayed rewards, where the only reward feedback is trajectory-wise ranking. SAIL bridges the advantages of IL and RL by interactively exploiting the demonstrations to catch up with the teacher and exploring the environment to yield demonstrations that surpass the teacher. Extensive empirical results show that not only does SAIL significantly improve the sample efficiency, but it also leads to higher asymptotic performance across different continuous control tasks, compared with the state-of-the-art.