Accelerating Reinforcement Learning using EEG-based implicit human feedback

Accelerating Reinforcement Learning using EEG-based implicit human feedback
复制标题

DOI:
10.1016/j.neucom.2021.06.064
复制
发表时间:
2021-10
期刊:
影响因子:
6
通讯作者:
Duo Xu;Mohit Agarwal;Ekansh Gupta;F. Fekri;R. Sivakumar
Duo Xu;Mohit Agarwal;Ekansh Gupta;F. Fekri;R. Sivakumar
中科院分区:
计算机科学2区
文献类型:
--
作者:
Duo Xu;Mohit Agarwal;Ekansh Gupta;F. Fekri;R. Sivakumar

文献摘要

相似文献

为强化学习(RL)代理提供人工反馈可以显着改善学习的各个方面。然而,先前的方法需要人类观察者明确地给出输入(例如,按钮,语音界面),在RL代理的学习过程中增加了人类的负担。此外,连续地提供明确的人类建议(反馈)并不总是可能的或限制性太大,例如,在这项工作中,我们研究通过错误相关电位(ErrP)形式的EEG捕获人类的内在反应作为隐式(和自然)反馈,为人类提供一种自然和直接的方法来改善RL Agent学习。因此,可以通过隐式反馈将人类智能与RL算法相结合,以加速RL Agent的学习。我们开发了三个相当复杂的2D离散导航游戏,实验评估所提出的工作的整体性能。主观实验也验证了使用ErrPs作为反馈的动机。我们工作的主要贡献如下,(i)我们提出并实验验证了ErrPs的零射击学习,其中ErrPs可以为一个游戏学习,并转移到其他看不见的游戏,(ii)我们提出了一种新的RL框架,用于通过ErrPs与RL代理集成隐式人类反馈,提高标签效率和对人类错误的鲁棒性,以及(iii)与以前的工作相比,我们将ErrPs的应用扩展到相当复杂的环境中,并通过真实的用户实验证明了我们的方法对加速学习的重要性。
Providing Reinforcement Learning (RL) agents with human feedback can dramatically improve various aspects of learning. However, previous methods require human observer to give inputs explicitly (e.g., press buttons, voice interface), burdening the human in the loop of RL agent’s learning process. Further, providing explicit human advise (feedback) continuously is not always possible or too restrictive, e.g., autonomous driving, disabled rehabilitation, etc. In this work, we investigate capturing human’s intrinsic reactions as implicit (and natural) feedback through EEG in the form of error-related potentials (ErrP), providing a natural and direct way for humans to improve the RL agent learning. As such, the human intelligence can be integrated via implicit feedback with RL algorithms to accelerate the learning of RL agent. We develop three reasonably complex 2D discrete navigational games to experimentally evaluate the overall performance of the proposed work. And the motivation of using ErrPs as feedbacks is also verified by subjective experiments. Major contributions of our work are as follows, (i) we propose and experimentally validate the zero-shot learning of ErrPs, where the ErrPs can be learned for one game, and transferred to other unseen games, (ii) we propose a novel RL framework for integrating implicit human feedbacks via ErrPs with RL agent, improving the label efficiency and robustness to human mistakes, and (iii) compared to prior works, we scale the application of ErrPs to reasonably complex environments, and demonstrate the significance of our approach for accelerated learning through real user experiments.