f-GAIL: Learning f-Divergence for Generative Adversarial Imitation Learning

f-GAIL: Learning f-Divergence for Generative Adversarial Imitation Learning
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Xin Zhang;Yanhua Li;Ziming Zhang;Zhi-Li Zhang
Xin Zhang;Yanhua Li;Ziming Zhang;Zhi-Li Zhang
中科院分区:
其他
文献类型:
--
作者:
Xin Zhang;Yanhua Li;Ziming Zhang;Zhi-Li Zhang

文献摘要

相似文献

模仿学习旨在从专家演示中学习一种策略,使学习者和专家行为之间的差异降至最低。为了量化这种差异,已经提出了不同的模仿学习算法,它们具有不同的预定偏差。这自然引发了这样一个问题:在给定一组专家演示的情况下,哪个分歧可以更准确地恢复专家政策,数据效率更高?在这项工作中,我们提出了一种新的生成性对抗性模仿学习(GAIL)模型$f$-Gail,该模型自动从$f$-发散族中学习差异度量以及能够产生专家样行为的策略。与各种预定义发散度量的IL基线相比,$f$-Gail在六个基于物理的控制任务中学习了更好的策略和更高的数据效率。
Imitation learning (IL) aims to learn a policy from expert demonstrations that minimizes the discrepancy between the learner and expert behaviors. Various imitation learning algorithms have been proposed with different pre-determined divergences to quantify the discrepancy. This naturally gives rise to the following question: Given a set of expert demonstrations, which divergence can recover the expert policy more accurately with higher data efficiency? In this work, we propose $f$-GAIL, a new generative adversarial imitation learning (GAIL) model, that automatically learns a discrepancy measure from the $f$-divergence family as well as a policy capable of producing expert-like behaviors. Compared with IL baselines with various predefined divergence measures, $f$-GAIL learns better policies with higher data efficiency in six physics-based control tasks.