Bootstrapping Human-Autonomy Collaborations by using Brain-Computer Interface of SSVEP for Multi-Agent Deep Reinforcement Learning

Bootstrapping Human-Autonomy Collaborations by using Brain-Computer Interface of SSVEP for Multi-Agent Deep Reinforcement Learning
复制标题

使用 SSVEP 脑机接口引导人类自主协作进行多智能体深度强化学习

DOI:
10.1109/ichms56717.2022.9980765
复制
发表时间:
2022
期刊:
2022 IEEE 3rd International Conference on Human-Machine Systems (ICHMS)
影响因子:
--
通讯作者:
Yi
Yi
中科院分区:
--
文献类型:
--
作者:
Joshua Ho;Chien;Chun;C. King;Chi;Tun;Yen;Yu;Yi

文献摘要

参考文献

被引文献

相似文献

由于复杂的机器设计的进步,人类自主团队(HAT)已经成为新兴的人工智能趋势之一,它允许与人类更密切的合作,同时作为人类最典型的助手执行道德、合理和适用的任务。基于HAT追求的集体目标和人与机器之间的权力共享,我们的研究旨在回答人类的脑机接口(BCI)是否有助于实现人类与强化学习(RL)代理的有效协作。它如何有效地促进人在环指导来引导智能体的训练?本研究提出了一个基于bci的系统,该系统与强化学习代理交互,作为人在环团队集成。脑机接口中稳态视觉诱发电位引发的神经反应促进了学习代理与人类的合作,并在游戏模拟环境中实现了这一目标。我们提出的系统NeuroRL的结果通过减少RL代理中开发和探索的非平稳性显示出显着的改进。使用bci辅助的human-in-the-loop,可以在早期调查中优化奖励,从而在训练中实现更有效的收敛。本研究提出的新设计可以扩展新兴HAT领域和基于知识的强化学习系统的发展,以适应动态环境中的各种应用。
Human-Autonomy Teaming (HAT) has become one of the emerging AI trends due to the advances in sophisticated machine design that allows closer cooperation with humans while performing moral, reasonable, and applicable tasks as humans’ most exemplary assistants. Based on HAT’s pursuing the collective goal and sharing the authority between humans and machines, our research aims at answering whether humans’ brain-computer interface (BCI) helps achieve efficient collaborations of human with Reinforcement Learning (RL) agents. How can it efficiently facilitate human-in-the-loop guidance to bootstrap the training of the agents? This study proposes a BCI-based system that interacts with RL agents as a human-in-the-loop teaming integration. The neural responses elicited by the Steady-State Visual Evoked Potential in BCI facilitate the collaboration of learning agents with humans and accomplish this goal in a game simulation environment. The results of our proposed system, NeuroRL, show significant improvement by reducing the non-stationarity of exploitations and explorations in the RL agents. With BCI-assisted human-in-the-loop, the rewards can be optimized during the early investigations to achieve more efficient convergence in the training. The novel design proposed in this study can extend the development of the emerging HAT field and knowledge-based RL systems for various applications in dynamic environments.
DOI: --
发表时间: 2015-11
期刊: CoRR
影响因子: --
作者:
T. Schaul;John Quan;Ioannis Antonoglou;David Silver
通讯作者: T. Schaul;John Quan;Ioannis Antonoglou;David Silver