Keeping Humans in the Loop: Teaching via Feedback in Continuous Action Space Environments

Keeping Humans in the Loop: Teaching via Feedback in Continuous Action Space Environments
复制标题

DOI:
10.1109/iros47612.2022.9982282
复制
发表时间:
2022-10
期刊:
2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Isaac S. Sheidlower;Allison Moore;Elaine Schaertl Short
Isaac S. Sheidlower;Allison Moore;Elaine Schaertl Short
中科院分区:
其他
文献类型:
--
作者:
Isaac S. Sheidlower;Allison Moore;Elaine Schaertl Short

文献摘要

被引文献

相似文献

交互式强化学习 (IntRL) 允许人类教师加速强化学习 (RL) 机器人的学习过程。然而,IntRL 在很大程度上仅限于具有离散动作空间的任务,其中动作相对较慢。这将 IntRL 的应用限制在更复杂和更具挑战性的机器人任务上,而现代 RL 特别适合这些任务。我们试图通过提出连续动作空间交互式强化学习(CAIR)来弥补这一差距:第一个连续动作空间 IntRL 算法,能够利用教师反馈在这些任务中超越最先进的 RL 算法。 CAIR 将从环境和教师中学到的策略结合到一个策略中,并根据这两个策略的协议按比例加权。这使得 CAIR 代理能够学习相对稳定的策略,尽管教师反馈可能存在噪音或粗糙。我们通过易于设计和理解的启发式预言老师在两个模拟机器人任务中验证了我们的方法。此外,我们通过 Amazon Mechanical Turk 在人类受试者研究中验证了我们的方法,并表明 CAIR 的性能优于交互式 RL 中的现有技术。
Interactive Reinforcement Learning (IntRL) allows human teachers to accelerate the learning process of Reinforcement Learning (RL) robots. However, IntRL has largely been limited to tasks with discrete-action spaces in which actions are relatively slow. This limits IntRL's application to more complicated and challenging robotic tasks, the very tasks that modern RL is particularly well-suited for. We seek to bridge this gap by presenting Continuous Action-space Interactive Reinforcement learning (CAIR): the first continuous action-space IntRL algorithm that is capable of using teacher feedback to out-perform state-of-the-art RL algorithms in those tasks. CAIR combines policies learned from the environment and the teacher into a single policy that proportionally weights the two policies based on their agreement. This allows a CAIR agent to learn a relatively stable policy despite potentially noisy or coarse teacher feedback. We validate our approach in two simulated robotics tasks with easy-to-design and - understand heuristic oracle teachers. Furthermore, we validate our approach in a human subjects study through Amazon Mechanical Turk and show CAIR out-performs the prior state-of-the-art in Interactive RL.