Active Reward Learning from Critiques

Active Reward Learning from Critiques
复制标题

DOI:
10.1109/icra.2018.8460854
复制
发表时间:
2018-05
期刊:
2018 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Yuchen Cui;S. Niekum
Yuchen Cui;S. Niekum
中科院分区:
其他
文献类型:
--
作者:
Yuchen Cui;S. Niekum

文献摘要

被引文献

相似文献

从示范算法中学习,例如逆增强学习,旨在为编程机器人提供一种自然机制,但通常需要过多的示范来捕捉任务的重要微妙之处。主动学习方法并没有盲目地要求盲目演示,而是利用不确定性来查询用户对具有高预期信息增益的州的动作标签。但是,这种方法仍然需要大量的标签来充分降低不确定性,并且可能不直觉,因为用户不习惯在单个脱节状态下确定最佳操作。为了解决这些缺点,我们提出了一种基于轨迹的新型主动贝叶斯逆增强学习算法,该算法是1)向用户查询对自动产生的轨迹的批评,而不是要求示范或动作标签,2)使用轨迹分段来加快批评 /批评 /批评 /标签过程和3)预测用户的批评是生成最有用的轨迹查询。我们评估了模拟域中的算法,发现它与先前的工作和随机基线相比。
Learning from demonstration algorithms, such as Inverse Reinforcement Learning, aim to provide a natural mechanism for programming robots, but can often require a prohibitive number of demonstrations to capture important subtleties of a task. Rather than requesting additional demonstrations blindly, active learning methods leverage uncertainty to query the user for action labels at states with high expected information gain. However, this approach can still require a large number of labels to adequately reduce uncertainty and may also be unintuitive, as users are not accustomed to determining optimal actions in a single out-of-context state. To address these shortcomings, we propose a novel trajectory-based active Bayesian inverse reinforcement learning algorithm that 1) queries the user for critiques of automatically generated trajectories, rather than asking for demonstrations or action labels, 2) utilizes trajectory segmentation to expedite the critique / labeling process, and 3) predicts the user's critiques to generate the most highly informative trajectory queries. We evaluated our algorithm in simulated domains, finding it to compare favorably to prior work and a randomized baseline.