Achieving Human-Robot Collaboration with Dynamic Goal Inference by Gradient Descent

Achieving Human-Robot Collaboration with Dynamic Goal Inference by Gradient Descent
复制标题

通过梯度下降动态目标推理实现人机协作

DOI:
10.1007/978-3-030-36711-4_49
复制
发表时间:
2019
期刊:
Neural Information Processing (Lecture Notes in Computer Science)
影响因子:
--
通讯作者:
Shigeki Sugano
Shigeki Sugano
中科院分区:
--
文献类型:
--
作者:
Shingo Murata;Wataru Masuda;Jiayi Chen;Hiroaki Arie;Tetsuya Ogata;Shigeki Sugano

文献摘要

相似文献

与人类合作伙伴的协作是智能机器人所期望的一项具有挑战性的任务。为了实现这一点,机器人需要能够与人类共享特定目标,并动态地推断人类是否改变了目标状态。在本文中,我们提出了一个基于神经网络的计算框架,具有基于梯度的目标状态优化,使机器人能够实现这种能力。该框架由卷积变分自编码器(ConvVAE)和具有长短期记忆(LSTM)架构的递归神经网络(RNN)组成,该架构学习将给定的目标图像映射到视觉预测。更具体地,视觉和目标特征状态首先由相应ConvVAE的编码器提取。然后,LSTM基于其当前状态生成视觉特征和运动预测,并根据提取的目标特征状态进行调节。在学习过程之后的协作期间,通过梯度下降来优化目标特征状态,以最小化预测和实际视觉特征状态之间的误差。这使得机器人能够仅从视觉观察动态地推断人类伙伴的情境(目标)变化。所提出的框架进行实验,涉及对象装配的人机协作任务进行评估。实验结果表明,机器人配备该框架可以与人类合作伙伴通过动态目标推理,即使在情况是模糊的。
Collaboration with a human partner is a challenging task expected of intelligent robots. To realize this, robots need the ability to share a particular goal with a human and dynamically infer whether the goal state is changed by the human. In this paper, we propose a neural network-based computational framework with a gradient-based optimization of the goal state that enables robots to achieve this ability. The proposed framework consists of convolutional variational autoencoders (ConvVAEs) and a recurrent neural network (RNN) with a long short-term memory (LSTM) architecture that learns to map a given goal image for collaboration to visuomotor predictions. More specifically, visual and goal feature states are first extracted by the encoder of the respective ConvVAEs. Visual feature and motor predictions are then generated by the LSTM based on their current state and are conditioned according to the extracted goal feature state. During collaboration after the learning process, the goal feature state is optimized by gradient descent to minimize errors between the predicted and actual visual feature states. This enables the robot to dynamically infer situational (goal) changes of the human partner from visual observations alone. The proposed framework is evaluated by conducting experiments on a human–robot collaboration task involving object assembly. Experimental results demonstrate that a robot equipped with the proposed framework can collaborate with a human partner through dynamic goal inference even when the situation is ambiguous.