Deep recurrent Q-learning of behavioral intervention delivery by a robot from demonstration data

Deep recurrent Q-learning of behavioral intervention delivery by a robot from demonstration data
复制标题

机器人根据演示数据进行行为干预的深度循环 Q 学习

DOI:
10.1109/roman.2017.8172429
复制
发表时间:
2017
期刊:
2017 26th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN)
影响因子:
--
通讯作者:
M. Begum
M. Begum
中科院分区:
--
文献类型:
--
作者:
Madison Clark;M. Begum

文献摘要

被引文献

相似文献

我们提出了一个从演示中学习(LfD)框架,该框架使用深度递归Q网络(DRQN)来学习如何从人类进行的演示中提供行为干预(BI)。经过训练的DRQN使机器人能够以自主方式提供类似的BI。BI是高度结构化的程序,其中患有发育迟缓/障碍(例如自闭症、ADHD等)的儿童接受过新行为和生活技能的训练。来自人机交互(HRI)研究的越来越多的轶事证据表明,BI受益于使用机器人作为交付工具。HRI对机器人干预的大部分研究依赖于遥控机器人。然而,对自主性的需求变得越来越明显,特别是在这些系统的实际部署方面。在基于机器人的BI中使用自主性的少数研究依赖于环境的精心挑选的特征,以触发正确的机器人动作。此外,这些自动化架构都没有尝试从人类演示中学习BI,尽管这似乎是最自然的学习方式。本文代表了第一次尝试设计一个机器人,使用LfD学习BI。我们生成一个模型,然后正确地预测适当的行动,准确率超过80%。据我们所知,这是第一次尝试在LfD框架内使用DRQN来学习嵌入在人类行为和行为中的高级推理。
We present a learning from demonstration (LfD) framework that uses a deep recurrent Q-network (DRQN) to learn how to deliver a behavioral intervention (BI) from demonstrations performed by a human. The trained DRQN enables a robot to deliver a similar BI in an autonomous manner. BIs are highly structured procedures wherein children with developmental delays/disorders (e.g. autism, ADHD, etc.) are trained to perform new behaviors and life-skills. Mounting anecdotal evidence from human-robot interaction (HRI) research has shown that BI benefits from the use of robots as a delivery tool. Most of the HRI research on robot-based intervention relies on tele-operated robots. However, the need for autonomy has become increasingly evident, especially when it comes to the real-world deployment of these systems. The few studies that have used autonomy in robot-based BI relied on hand-picked features of the environment in order to trigger correct robot actions. Additionally, none of these automated architectures attempted to learn the BI from human demonstrations, though this appears to be the most natural way of learning. This paper represents the first attempt to design a robot that uses LfD to learn BI. We generate a model then correctly predict appropriate actions with greater than 80% accuracy. To the best of our knowledge, this is the first attempt to employ DRQN within an LfD framework to learn high level reasoning embedded in human actions and behaviors simply from observations.