Goal-Conditioned Variational Autoencoder Trajectory Primitives with Continuous and Discrete Latent Codes

Goal-Conditioned Variational Autoencoder Trajectory Primitives with Continuous and Discrete Latent Codes
复制标题

具有连续和离散潜在代码的目标条件变分自动编码器轨迹原语

DOI:
10.1007/s42979-020-00324-7
复制
发表时间:
2020
期刊:
SN Computer Science
影响因子:
--
通讯作者:
Shuhei Ikemoto
Shuhei Ikemoto
中科院分区:
--
文献类型:
--
作者:
Takayuki Osa;Shuhei Ikemoto

文献摘要

参考文献

相似文献

模仿学习是向机器人系统教授动作的一种直观方法。虽然以前的研究提出了各种方法来建模演示的运动基元,但现有方法的局限性之一是轨迹的形状是在高维空间中编码的。轨迹表示的高维可能是后续过程中的瓶颈,例如规划基元运动序列。我们通过学习机器人轨迹的潜在空间来解决这个问题。如果能够学习轨迹的潜变量,即使用户不是专家,也可以用它以直观的方式调整轨迹。我们提出了一个用学习低维潜在空间的神经网络来建模演示轨迹的框架。我们的神经网络结构建立在具有离散和连续潜变量的变分自动编码器(VAE)的基础上。我们对已有的VAE结构进行了扩展,得到了以轨迹的目标位置为条件的解码器,将其推广到不同的目标位置。虽然VAE进行的推理并不准确,但通过将投影合并到解空间上,可以将广义目标位置的定位误差降低到1 mm以下。为了应对海量训练数据的需求,我们采用了一种受计算机视觉领域中常用的数据增强技术启发的轨迹增强技术。在提出的框架中,编码多种轨迹的潜在变量以无监督的方式学习,尽管现有方法通常需要标签信息来建模不同的行为。学习的解码器可以用作运动规划器,其中用户可以通过设置潜变量来指定目标位置和轨迹类型。实验结果表明,我们的神经网络可以使用有限数量的演示轨迹来训练,并且可以学习可解释的潜在表示。
Imitation learning is an intuitive approach for teaching motion to robotic systems. Although previous studies have proposed various methods to model demonstrated movement primitives, one of the limitations of existing methods is that the shape of the trajectories is encoded in high dimensional space. The high dimensionality of the trajectory representation can be a bottleneck in the subsequent process such as planning a sequence of primitive motions. We address this problem by learning the latent space of the robot trajectory. If the latent variable of the trajectories can be learned, it can be used to tune the trajectory in an intuitive manner even when the user is not an expert. We propose a framework for modeling demonstrated trajectories with a neural network that learns the low-dimensional latent space. Our neural network structure is built on the variational autoencoder (VAE) with discrete and continuous latent variables. We extend the structure of the existing VAE to obtain the decoder that is conditioned on the goal position of the trajectory for generalization to different goal positions. Although the inference performed by VAE is not accurate, the positioning error at the generalized goal position can be reduced to less than 1 mm by incorporating the projection onto the solution space. To cope with requirement of the massive training data, we use a trajectory augmentation technique inspired by the data augmentation commonly used in the computer vision community. In the proposed framework, the latent variables that encodes the multiple types of trajectories are learned in an unsupervised manner, although existing methods usually require label information to model diverse behaviors. The learned decoder can be used as a motion planner in which the user can specify the goal position and the trajectory types by setting the latent variables. The experimental results show that our neural network can be trained using a limited number of demonstrated trajectories and that the interpretable latent representations can be learned.
通过优化的运动原语
DOI: 10.1109/icra.2015.7139510
发表时间: 2015
期刊: 2015 IEEE International Conference on Robotics and Automation (ICRA)
影响因子: --
作者:
A. Dragan;Katharina Muelling;J. Bagnell;S. Srinivasa;S. Srinivasa
通讯作者: S. Srinivasa
DOI: 10.1007/978-3-319-58347-1_10
发表时间: 2017-01-01
期刊: DOMAIN ADAPTATION IN COMPUTER VISION APPLICATIONS
影响因子: --
作者:
Ganin, Yaroslav;Ustinova, Evgeniya;Lempitsky, Victor
通讯作者: Lempitsky, Victor
DOI: 10.1561/2300000053
发表时间: 2018-03
期刊: ArXiv
影响因子: --
作者:
Takayuki Osa;J. Pajarinen;G. Neumann;J. Bagnell;P. Abbeel;Jan Peters
通讯作者: Takayuki Osa;J. Pajarinen;G. Neumann;J. Bagnell;P. Abbeel;Jan Peters
DOI: 10.3389/fnbot.2019.00022
发表时间: 2019-05-31
影响因子: 3.1
作者:
Arnold, Solvi;Yamazaki, Kimitoshi
通讯作者: Yamazaki, Kimitoshi