Policy Gradient Bayesian Robust Optimization for Imitation Learning

Policy Gradient Bayesian Robust Optimization for Imitation Learning
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Zaynah Javed;Daniel S. Brown;Satvik Sharma;Jerry Zhu;A. Balakrishna;Marek Petrik;A. Dragan;Ken Goldberg
Zaynah Javed;Daniel S. Brown;Satvik Sharma;Jerry Zhu;A. Balakrishna;Marek Petrik;A. Dragan;Ken Goldberg
中科院分区:
其他
文献类型:
--
作者:
Zaynah Javed;Daniel S. Brown;Satvik Sharma;Jerry Zhu;A. Balakrishna;Marek Petrik;A. Dragan;Ken Goldberg

文献摘要

被引文献

相似文献

由于很难为许多现实世界的问题指定奖励,人们越来越关注从人类反馈中学习奖励,比如示范。然而,通常有许多不同的奖励函数来解释人类的反馈,这给代理留下了真正的奖励函数是什么的不确定性。虽然大多数策略优化方法通过优化预期性能来处理这种不确定性,但许多应用程序需要规避风险的行为。我们提出了一种新的策略梯度式稳健优化方法PG-BROIL,它优化了一个软鲁棒目标,平衡了预期性能和风险。据我们所知,PG-BROIL是第一个对报酬假设分布稳健的策略优化算法,可以扩展到连续的MDP。结果表明,PG-BROIL可以产生从风险中性到风险厌恶的一系列行为,并且在通过对冲不确定性而不是寻求唯一地识别演示者的奖励函数来从模棱两可的演示中学习时,其性能优于最新的模仿学习算法。
The difficulty in specifying rewards for many real-world problems has led to an increased focus on learning rewards from human feedback, such as demonstrations. However, there are often many different reward functions that explain the human feedback, leaving agents with uncertainty over what the true reward function is. While most policy optimization approaches handle this uncertainty by optimizing for expected performance, many applications demand risk-averse behavior. We derive a novel policy gradient-style robust optimization approach, PG-BROIL, that optimizes a soft-robust objective that balances expected performance and risk. To the best of our knowledge, PG-BROIL is the first policy optimization algorithm robust to a distribution of reward hypotheses which can scale to continuous MDPs. Results suggest that PG-BROIL can produce a family of behaviors ranging from risk-neutral to risk-averse and outperforms state-of-the-art imitation learning algorithms when learning from ambiguous demonstrations by hedging against uncertainty, rather than seeking to uniquely identify the demonstrator's reward function.